For What Reason? Interpreting Models' Encoding of Causation and Antithesis
This paper investigates how instruction-tuned Transformer models encode causation and antithesis discourse relations, revealing that predictive decisions are made at different stages across layers and that these models exhibit asymmetric representations of discourse-based reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand a conversation not just by hearing the words, but by feeling the invisible threads that tie them together. In the world of artificial intelligence, these threads are called "discourse relations." Think of them as the glue that holds a story together, telling you if one sentence is the reason for the next, or if it's a surprise that contradicts it. For a computer to be truly helpful and safe, it needs to understand these connections, not just the dictionary definitions of the words. If a robot can't tell the difference between "I studied, so I passed" (a happy cause-and-effect) and "I studied, yet I failed" (a sad contradiction), it might give you terrible advice or misunderstand your feelings. This paper dives into the "brain" of modern AI to see how it actually processes these tricky relationships.
The researchers, Abhidip Bhattacharyya and Shira Wein, decided to play detective inside two popular AI models (LLaMA and Mistral) to figure out how they distinguish between causation (things happening because of something else) and antithesis (things happening in spite of something else). They set up a simple game: show the AI a sentence with a blank, like "John studied hard, so he was ___ to pass," and ask it to fill in the blank with either "able" or "unable." By swapping the connecting words (changing "so" to "yet") or flipping the meaning of the verbs, they could see exactly how the AI's internal gears turned to make the right choice.
Here is what they found, and it's a bit like discovering how a factory assembly line works.
The Assembly Line of Thought
The team discovered that the AI doesn't make its decision all at once. Instead, it's a relay race. When the AI reads the clue words like "so" or "yet," the early and middle layers of its brain (think of these as the first few floors of a skyscraper) immediately start figuring out the local meaning. They are the ones saying, "Oh, 'so' usually means a good outcome, while 'yet' means a twist!"
However, these early layers don't have the final say. As the information travels up the tower to the higher layers, the decision gets locked in. The researchers found that once the middle layers have done their job, the top layers mostly just carry that decision forward to the finish line. If you try to mess with the top layers by swapping their "thoughts" with a different sentence, the AI doesn't change its mind—it's too late; the decision is already made. But if you mess with the middle layers right when the AI sees the word "so" or "yet," you can actually flip its answer. This suggests the middle layers are the real decision-makers for these specific clues.
The Secret Bias
One of the most interesting findings is that the AI isn't perfectly neutral. The researchers found that certain layers have a built-in bias. Some layers seem to have a strong preference for the answer "unable," while others lean heavily toward "able," regardless of what the sentence actually says. It's like having a team of judges where some are naturally grumpy and want to say "no," while others are optimistic and want to say "yes." The final answer is a tug-of-war between these biased layers until the truth wins out.
The Role of the Verb
The study also showed that the AI is smart enough to look beyond just the connecting words. If the sentence uses a negative verb (like "breaks rules" instead of "studied hard"), the AI has to work harder to figure out if "able" or "unable" fits. The researchers found that the AI uses different internal pathways depending on whether the verb is positive or negative. When the verbs were negative, the connections inside the AI's brain were less stable and more chaotic, suggesting that negative emotions or actions make the AI's reasoning process a bit wobblier.
The Bottom Line
In short, this paper suggests that AI models don't just "know" the answer; they build it step-by-step. The early parts of the brain spot the clues, the middle parts make the call, and the top parts just deliver the verdict. While the AI is very good at this task (getting it right almost all the time), it relies on specific, sometimes biased, internal circuits to do it. This helps us understand that for AI to be truly aligned with human values, we need to know not just what it answers, but how it decides, because the path to the answer is just as important as the answer itself.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.