Three Lenses on Linguistic Knowledge in Pretrained Transformers: A Phenomenon-Centered Review
This review synthesizes evidence from representational, behavioral, and mechanistic perspectives on subject-verb agreement, negation, and compositional generalization in pretrained Transformers to reveal distinct patterns of convergence and fragmentation, ultimately proposing an accessibility-and-use diagnostic to clarify current limitations and guide future cross-level research.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language is a system of rules that allows us to take simple ideas and combine them into infinite new thoughts. For decades, scientists have wondered if the massive computer programs that now write fluent text and answer complex questions have truly learned these rules, or if they are merely mimicking patterns they have seen before. These programs, known as pretrained language models, are trained on vast amounts of human writing. They can produce sentences that sound perfectly natural, leading many to assume they understand grammar and meaning in the same way humans do. However, producing the right words is not the same as knowing why those words belong together. A program might get a sentence right by accident, or by relying on a shortcut, without actually grasping the underlying logic. To solve this mystery, researchers have developed three different ways to look inside these models: checking what information is stored in their internal memory, testing what they actually produce in their output, and tracing the specific pathways that cause them to make a decision.
A new review by researchers Yizhe Wang and Zhenhua Ling brings these three ways of looking together to see if they tell the same story. The authors focused on three specific areas of language that are essential for understanding how these models work: how verbs match with their subjects, how the word "not" changes the meaning of a sentence, and how the models combine familiar parts to create new, complex ideas. By gathering and comparing hundreds of studies that used these three different methods, the researchers found that the answer depends entirely on which part of language you are asking about. Sometimes the evidence from all three methods lines up perfectly; other times, the methods tell conflicting stories, revealing gaps in what the models actually know.
The clearest success story the researchers found involves subject-verb agreement, the rule that a singular subject needs a singular verb, like "the cat runs," while a plural subject needs a plural verb, like "the cats run." In this area, the three different ways of looking at the models agree. When researchers looked inside the model's internal states, they could find a clear signal for whether a noun was singular or plural. When they tested the model's output, it almost always chose the correct verb form. Most importantly, when researchers intervened to change that internal signal, the model's behavior changed in a predictable way, forcing it to pick the wrong verb. This convergence suggests that for simple grammatical rules, these models have genuinely learned the relationship and use it to make decisions.
The picture becomes much more complicated when looking at negation, or how the models handle the word "not." Here, the three lenses do not agree. The internal checks show that the models can detect the presence of the word "not" and understand roughly where it applies in a sentence. However, when the models are asked to produce text or answer questions based on a negated sentence, they often fail to apply the correct meaning. They might see the word "not" but still act as if the sentence were positive. The researchers found that while the models can spot the marker, they do not reliably use it to flip the truth of the statement. It is as if the model sees the sign but does not follow the instruction it gives. This mismatch means that while the models have the pieces of information, they are not consistently using them to understand the full meaning.
The third area, which involves combining different parts of language to create new meanings, shows the most confusion. This is the ability to take known words and rules and apply them to situations the model has never seen before. The evidence here is fragmented. Some studies show that the models can handle simple combinations, especially when researchers provide hints or break the task down into smaller steps. Other studies show that when the models are left to figure out complex chains of reasoning on their own, they often fail. The problem is that the different studies are not looking at the same thing. Some are testing if the model can recognize a pattern, while others are testing if it can build a new idea from scratch. Because the researchers could not find a single, consistent way to measure this ability across all studies, they could not conclude whether the models truly possess a general skill for combining ideas or if they are just good at specific, narrow tricks.
The authors of the review suggest that these different outcomes happen because some language features are easier to isolate and use than others. For subject-verb agreement, the feature is simple and isolated, so the model can find it and use it. For negation, the model can find the marker but struggles to use it to change the whole meaning. For complex combinations, the feature is so spread out and difficult to define that the model cannot even isolate it clearly. The researchers conclude that we cannot assume these models know language in a general sense. Instead, their knowledge is specific to the task. They are excellent at some things, like matching verbs, but they are inconsistent at others, like understanding negation, and they are still unproven at the most complex task of creating new ideas from old parts. This finding changes how we should evaluate these systems: we cannot just look at whether they get the right answer, but we must also check if they are using the right logic to get there.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.