How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines
This paper introduces new Cantonese ParGram resources and presents a controlled evaluation showing that while advanced LLMs can generate locally plausible grammar components, they struggle with complex formal constraints, indicating that human linguistic expertise remains essential for grammar engineering.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language is a living, breathing thing, but to teach a computer to understand it, linguists often have to build a rigid skeleton first. This field, known as grammar engineering, involves writing detailed rulebooks that describe exactly how words fit together to form meaning. Think of these rulebooks not as simple lists of correct sentences, but as complex instruction manuals that tell a machine how to break a sentence down into its functional parts, identifying who is doing what to whom, and how those roles connect across different languages. For decades, creating these manuals has been a slow, painstaking process requiring deep human expertise, as engineers must manually encode the subtle logic of every sentence structure. Recently, a new tool has entered the workshop: large language models, the same powerful artificial intelligence systems that can write essays or answer questions. The big question for linguists is whether these AI systems can actually help build the complex rulebooks needed to teach computers about language, or if they are simply too prone to error for such precise work.
A researcher set out to test this question using Cantonese, a widely spoken language that has historically lacked the detailed digital resources available for English. They created a new set of high-quality reference materials, essentially a gold standard of correct sentence structures and their underlying logical relationships, to serve as a benchmark. With this foundation in place, they ran a series of controlled experiments to see if two different AI models could generate the necessary grammar rules from scratch. The researcher did not ask the AI to simply translate text; instead, they asked it to act as a grammar engineer, producing the specific technical components required to make a computer parse sentences correctly. They tested the models under different conditions: sometimes giving them just a raw sentence, and other times giving them the raw sentence along with the correct logical structure they were supposed to aim for. They also varied the difficulty, asking the AI to handle single sentence types versus groups of related sentences that interact in complex ways.
The results revealed a clear hierarchy of capability. The more advanced AI model consistently outperformed the other, producing grammar rules that were far more likely to be correct and useful. However, the method of prompting mattered just as much as the model itself. When the researcher provided the AI with the target logical structure alongside the sentence, the resulting grammar rules were significantly better than when the AI had to guess the structure from the sentence alone. This suggests that while the AI can recognize patterns, it struggles to deduce the deep, invisible logic of language without a clear guide. Even the best-performing model, when faced with the challenge of handling multiple different sentence types at once, began to falter. It could often generate rules that looked correct on the surface, but failed when those rules had to work together to analyze a complex sentence. The AI frequently missed the subtle connections between different parts of the grammar, such as how a specific word in one part of a sentence must logically match a specific role in another part.
The study concludes that these AI systems are not yet ready to replace human linguists in building these complex language systems. They cannot independently create a robust, working grammar from a list of sentences. Instead, their most useful role appears to be as a supportive assistant. The research suggests a workflow where the AI might generate a first draft of the logical structures or rule candidates, which a human expert then reviews, corrects, and refines. In this partnership, the AI handles the heavy lifting of generating possibilities, while the human provides the critical judgment needed to ensure the final system is logically sound and accurate. The findings confirm that while artificial intelligence can accelerate parts of the process, the deep, structural understanding required to build a reliable grammar engine remains firmly in the realm of human expertise.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.