The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models
This paper investigates whether language models can sustain philosophical conceptual analysis through iterated counterexample and repair cycles, finding that while they can engage in the process, the loop quickly yields diminishing returns characterized by increasing verbosity without improved accuracy and a tendency for the models to over-accept invalid counterexamples compared to human judges.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a group of philosophers playing a high-stakes game of "Gotcha."
One person proposes a simple definition for a word, like "A game is something you do for fun." The next person tries to break that definition by finding a real-life situation where the definition fails (a counterexample). For instance, "What about a professional poker player who is stressed and hates the game but keeps playing to win money? That's still a game, but it's not 'for fun'."
The first person then has to fix their definition to include the poker player but still exclude things that aren't games. They try again, and the cycle repeats. This is called Conceptual Analysis, and it's how philosophers have refined ideas for thousands of years.
The researchers in this paper asked: Can AI language models play this game?
They set up a digital version of this game where one AI acts as the "Proposer" (making definitions) and another acts as the "Critic" (finding holes in the definitions). They let them play this loop over and over again—sometimes 50 times in a row—to see if the AI gets smarter at defining things like "friend," "lie," or "sandwich."
Here is what they found, translated into everyday terms:
1. The AI Judge is a bit too generous
In the game, the "Critic" needs to spot bad definitions. The researchers had human experts and an AI judge rate the AI's counterexamples.
- The Result: The AI judge said "Yes, that's a valid counterexample!" about twice as often as the human experts did.
- The Metaphor: Imagine a strict teacher (the human) and a very eager student (the AI) grading a test. The student thinks every wrong answer is a brilliant trick question, while the teacher knows some are just mistakes. However, they did agree on the "big hits"—the really obvious flaws. So, while the AI is a bit too easy on itself, it isn't completely hallucinating.
2. The "Longer is Better" Trap
The researchers expected that as the AI played more rounds, the definitions would get sharper and more accurate, just like a human philosopher refining an idea over years.
- The Result: Instead of getting better, the definitions just got longer and more wordy.
- The Metaphor: Imagine trying to describe a "dog" to an alien.
- Round 1: "A dog is a furry animal." (Too broad: includes cats).
- Round 2: "A dog is a furry animal that barks." (Too broad: includes some seals).
- Round 50: "A dog is a furry animal that barks, has four legs, is not a seal, is not a cat, usually has a tail, likes bones, and is not a robot, unless it's a specific type of robot dog..."
The AI kept adding "exceptions" and "rules" to patch every hole, but it never actually figured out the core truth of what a dog is. The definition became a massive, clunky laundry list of rules rather than a clear, simple insight.
3. Some Concepts are Just Harder to Pin Down
The researchers tested 20 different concepts. They found that some words were easy for the AI to "break" with counterexamples, while others were nearly impossible.
- Easy to break: Words like "neighbor" or "friend." The AI could easily find weird edge cases (e.g., "Is a neighbor someone you've never met but lives next door?").
- Hard to break: Words like "mistake" or "expert." These concepts seemed to resist the game. The AI struggled to find valid counterexamples, suggesting these ideas are inherently fuzzier or harder to define with strict rules.
4. The "Memory" Didn't Help
The researchers tried two versions of the game:
- Memoryless: The AI only sees the current definition.
- With History: The AI sees the entire history of every previous definition and counterexample.
- The Result: Having the history didn't make the AI any better. It didn't learn from its past mistakes. It just kept making the same kind of mistakes, but with more words.
The Bottom Line
The paper concludes that while AI can play the "Counterexample Game" and understand the rules, it hits a wall very quickly. It doesn't get smarter with practice; it just gets more verbose.
The researchers suggest this is a great way to test AI. If an AI can't sustain high-level, iterative philosophical reasoning (getting better over time), it tells us something important about how these models think: they are good at patching holes in the short term, but they struggle to build a stable, deep understanding of complex ideas over the long term.
In short: The AI can play the game, but it's like a player who keeps adding more and more rules to the rulebook without ever actually understanding the sport.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.