← Latest papers
💬 NLP

Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages

By analyzing 29 languages across five model families, this study demonstrates that multilingual large language models utilize partially shared computational circuits for subject-verb agreement, with the degree of cross-lingual overlap increasing in languages that overtly express morphological inflection.

Original authors: Isabella Gidi, Antonio Almudévar, Core Francisco Park, Naomi Saphra, Ricard Marxer

Published 2026-08-20
📖 4 min read☕ Coffee break read

Original authors: Isabella Gidi, Antonio Almudévar, Core Francisco Park, Naomi Saphra, Ricard Marxer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Modern computers that read and write in many languages have become remarkably skilled, often translating or answering questions in a dozen tongues with a single brain. Yet, for all their fluency, scientists do not fully understand how these systems work inside. A central mystery is whether a multilingual computer uses one universal set of rules to handle grammar for every language, or if it builds a separate, unique engine for each one. This question matters because if the machine uses the same internal machinery for different languages, then insights gained from studying one language could help us understand all of them. If it builds separate engines, however, then what we learn about English might tell us nothing about how it handles Spanish or Chinese. The specific puzzle this research tackles is how these models handle subject-verb agreement, the grammatical rule that forces a verb to change its form to match the person doing the action, such as changing "walk" to "walks" when the subject is "he."

To solve this, researchers examined how five different families of open-source language models process this grammatical rule across twenty-nine languages. They focused on a specific moment in the computer's thinking process: the split second before it predicts the final letter or word of a verb. By using a technique that swaps the internal signals from a correct sentence into a slightly altered, incorrect one, the team could pinpoint exactly which parts of the computer's network were responsible for getting the grammar right. They tested this method on languages that change their verbs heavily, like Spanish, languages that do not change them at all, like Chinese, and English, which changes them only in a few specific cases.

The investigation revealed that these models do not rely on entirely separate engines for every language, nor do they use a single, identical engine for all. Instead, they appear to reuse a shared set of internal pathways, but only when the task demands it. When the researchers looked at languages where the verb must visibly change to match the subject, they found that the same specific parts of the computer's network were active across almost all of them. These shared pathways were most consistent when the computer had to recover the exact difference between two verb forms, such as distinguishing between "I walk" and "he walks." In these moments, the models seemed to rely on a common, shared circuitry to perform the grammatical calculation.

However, this sharing is not uniform. When the researchers looked at languages where the verb does not change at all, the internal pathways looked very different. The computer did not seem to use the same shared machinery for these languages. Even more telling was the behavior of English. Because English only changes its verbs in the third-person singular, the computer's internal wiring looked different from the heavily conjugating languages most of the time. But the moment the computer had to produce that specific "he walks" form, its internal activity shifted to look almost exactly like the patterns seen in languages that change their verbs constantly. This suggests that the computer does not decide to use a shared engine based on the language itself, but rather based on the specific grammatical operation it is performing at that moment.

The study also examined how these shared parts of the network actually behave. They found that the specific components responsible for agreement in different languages pay attention to the same parts of the sentence. When a model is deciding on a verb form, these shared components focus their attention on the subject pronoun and the verb itself, regardless of whether the language is Spanish, French, or the specific case in English. This indicates that the overlap is not just a coincidence of location; the shared parts are actually performing the same job of routing information about the subject to the verb.

Ultimately, the findings suggest that multilingual computers are not simply memorizing separate rules for every language, nor are they using a single, rigid system for everything. They appear to have a flexible architecture that reuses shared internal structures whenever a language requires the same grammatical operation to be performed visibly. The degree to which these structures are shared depends less on the identity of the language and more on whether the language forces the computer to explicitly calculate and display that grammatical difference. This provides a clearer picture of how artificial intelligence learns to handle the complex, varied rules of human grammar, showing that it builds its understanding by reusing functional tools rather than constructing entirely new ones for every tongue it learns.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →