Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks
This paper details the development of Cantonese and Irish treebanks within the Parallel Grammar Project while evaluating the utility of multilingual LLMs in grammar engineering, finding that while these models offer limited value for suggesting alternative analyses, they currently lack the cross-linguistic abstraction capabilities necessary to replace expert-driven linguistic verification.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine trying to build a universal translator for the human mind, not just for words, but for the deep, invisible rules that make sentences work. This is the world of grammar engineering. Think of it like being an architect who doesn't just draw a house, but writes the actual code that tells a robot how to build it, brick by brick, ensuring the roof doesn't fall in. To do this, linguists use special maps called treebanks. If a sentence is a building, a treebank is the blueprint that shows exactly how every word connects to every other word, capturing who did what to whom. But here's the tricky part: when you try to build these blueprints for two very different languages—like one that flows like a river and another that jumps like a frog—you need a master plan that keeps the core structure the same while letting the surface details differ. This is the challenge of parallel grammar: making sure the "soul" of the sentence is identical across languages, even if the "clothes" it wears are totally different.
Now, enter the new kid on the block: Large Language Models (LLMs). These are the super-smart AI chatbots that have read almost everything on the internet. They are great at guessing the next word in a sentence, but can they actually understand the deep architectural rules of grammar? Can they help the human architects build these blueprints faster, or are they just fancy guessers that might accidentally build a house with a door in the ceiling? This paper asks exactly that question, using two very distinct languages: Cantonese (a language from Southern China with very few word endings) and Irish (a Celtic language with lots of word endings and a different word order). The researchers wanted to see if an AI could translate between these two and draw the correct grammatical blueprints for them, acting as a helpful assistant to human experts.
The Experiment: AI as a Grammar Apprentice
The researchers set up a test drive with a specific AI model (OpenAI's gpt-oss-120b) and gave it a stack of 50 sentences. These sentences were like a "driver's license test" for grammar, covering everything from simple statements to complex sentences with relative clauses and tricky verb structures. They asked the AI to do two main jobs:
- Translate: Turn a sentence from Cantonese into Irish, or vice versa.
- Draw the Blueprint: Create a formal grammatical map (called an f-structure) that shows the deep relationships between words, ignoring the surface order.
They tried this in different ways: asking the AI in English, asking it in the target language, and asking it to do both translation and drawing at the same time.
The Results: A Mixed Bag of Brilliance and Confusion
The findings were a bit like watching a talented but confused apprentice. The AI showed it could sometimes see the big picture, but it often stumbled on the details.
The Translation Trouble
When it came to translating, the AI struggled mightily. It was like a student who knows the vocabulary but keeps mixing up the dialects.
- The "Mandarin" Mix-up: When asked to speak Cantonese, the AI often slipped into Mandarin. For example, it used the word hǎo bàng (great) instead of the correct Cantonese hóu lèk. It's like asking someone to speak with a Scottish accent, and they accidentally use American slang instead.
- The "Invented" Words: When translating into Irish, the AI didn't just make mistakes; it made up words. The word for "tractor" (tarracóir) was sometimes turned into "tarpaulin," "robber," "rabbit," or even "thief." It seemed to be guessing based on how the word sounded or what context it was in, rather than knowing the actual meaning.
- The Success Rate: The numbers were sobering. When asked to translate from Irish to Cantonese, only about 22% to 24% of the translations were both accurate in meaning and natural-sounding. For Cantonese to Irish, it was even lower, with only 6% to 8% of translations hitting the mark. Interestingly, changing the language of the instructions (the "prompt") didn't help at all; the AI performed just as poorly whether asked in English or the target language.
The Blueprint Drawing
When it came to drawing the grammatical blueprints (the f-structures), the AI did slightly better, but still had a long way to go.
- English was Easy: On English sentences, the AI got it right (rated "Excellent" or "Good") about 46% of the time.
- Cantonese was Okay: For Cantonese, it managed a correct structural analysis about 34% of the time.
- Irish was Hard: For Irish, the success rate dropped to just 20%. The researchers suspect this is because Irish has more complex word endings and a different sentence order (Verb-Subject-Object) that the AI wasn't as familiar with.
- The "Shared" Blueprint Challenge: The hardest task was asking the AI to find the shared deep structure between Cantonese and Irish. Here, the AI mostly failed, with 44% of the Cantonese-to-Irish attempts and 38% of the Irish-to-Cantonese attempts being rated as "Bad." It struggled to abstract the common rules while ignoring the surface differences.
What the AI Got Wrong (and Right)
The AI wasn't entirely useless. It did manage to capture the basic "who did what" relationships in about 70% of the cases, even if the details were messy. It sometimes suggested interesting alternative ways to look at a sentence, which could spark an idea for a human linguist.
However, the errors were revealing.
- Stereotypes: The AI assumed that "farmers" and "drivers" were always male, assigning them masculine gender tags even though neither Cantonese nor Irish grammatically requires gender for these words. It was relying on its own "world knowledge" rather than the rules of the language.
- Confusing Concepts: It often mixed up different types of sentence structures, treating a relative clause (like "the car that I bought") as if it were a direct object, showing it didn't fully grasp the theoretical differences.
- Hallucinations: The AI invented words and meanings, like translating "bathed in the river" as "drowned himself in the river," completely changing the story.
The Bottom Line
The paper concludes that while these AI models are powerful tools, they are not ready to replace human experts in building complex grammars. They can offer a "rough draft" or a hint of an idea, but they cannot be trusted to get the details right. The AI is like a very fast, very confident student who has read a lot of books but hasn't actually practiced the craft enough to build a sturdy house.
The researchers emphasize that human expertise is still essential. The AI's outputs need careful checking and verification by people who truly understand the languages. The study suggests that until AI models are trained on much more specific, high-quality data for these "under-resourced" languages, they will remain more of a curiosity than a reliable partner in the serious work of grammar engineering. The dream of a fully automated grammar builder is still a ways off, but this experiment helped map out exactly where the current technology trips over its own shoelaces.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.