← Latest papers
💬 NLP

Impact of Task Phrasing on Presumptions in Large Language Models

This study demonstrates that task phrasing significantly influences the formation of presumptions in large language models, affecting their adaptability and logical reasoning in decision-making tasks like the iterated prisoner's dilemma, thereby highlighting the critical need for neutral phrasing to ensure safety and reliability.

Original authors: Kenneth J. K. Ong

Published 2026-05-04
📖 5 min read🧠 Deep dive

Original authors: Kenneth J. K. Ong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a very smart, well-read robot how to play a game. This robot has read millions of books, articles, and stories, so it knows the "rules of the world" very well. But here's the catch: sometimes, the robot is so good at remembering what it usually sees that it stops looking at what is actually in front of it.

This paper is like a detective story investigating exactly that problem. The author, Kenneth Ong, wanted to see if Large Language Models (LLMs)—the brains behind AI chatbots—get stuck in their own assumptions when the rules of a game change.

The Game: A Twist on "Prisoner's Dilemma"

To test this, the author used a classic game theory scenario called the Iterated Prisoner's Dilemma. Think of this as a game between two suspects in a police station.

  • The Normal Rules: If both stay quiet (Cooperate), they get a short sentence. If one betrays the other (Defects), the betrayer goes free, and the quiet one gets a long sentence.
  • The Trap: The author created a "flipped" version of the game. In this version, the rewards are swapped. Now, betraying the other person actually leads to a worse outcome for you, while staying quiet is the best move.

The goal was simple: If the AI is truly logical, it should look at the new rules and change its strategy. If it is just relying on "presumptions" (what it thinks the game should be), it will keep playing the old way, even though the rules have changed.

The Three Experiments: How the Robot Was Asked to Play

The author ran three different tests, changing how the instructions were written each time.

1. The "Name Drop" Test (Explicit Naming)

  • The Setup: The prompt explicitly said, "You are playing the Iterated Prisoner's Dilemma." It used the words "prisoner," "cooperate," and "defect."
  • The Result: The AI failed completely. Even when the rules were flipped, the AI kept choosing to "cooperate" (or "defect," depending on the model) exactly as it would in the original game.
  • The Analogy: It's like telling a chef, "Make me a classic Beef Wellington," but then handing them a recipe for a vegetarian stew. The chef, hearing "Beef Wellington," ignores the new recipe and starts cooking beef anyway because that's what the name implies. The AI was so focused on the famous name of the game that it ignored the actual numbers on the page.

2. The "Vague Description" Test (Hidden Name)

  • The Setup: The prompt removed the words "Prisoner's Dilemma" but still used the words "prisoner," "cooperate," and "defect."
  • The Result: The AI still struggled. It seemed to figure out it was a "Prisoner's Dilemma" just by the context clues (the jail sentences and the specific terms). It still relied on its training data rather than the new, flipped rules.
  • The Analogy: This is like describing a "Beef Wellington" without using the name, but saying, "It's a famous British dish with puff pastry and beef." The chef still knows what you want and ignores the new instructions.

3. The "Neutral Code" Test (Abstract Variables)

  • The Setup: The prompt removed everything familiar. No "prisoners," no "cooperate," no "defect." Instead, it used abstract labels: "Choice X" and "Choice Y," and "losing dollars" instead of "jail time."
  • The Result: Success! When the AI couldn't rely on its memory of the famous game, it actually looked at the new rules. When the rules were flipped, the AI changed its strategy to match the new math.
  • The Analogy: This is like giving the chef a list of ingredients and instructions without naming the dish. "Mix flour, butter, and beef." If you change the instructions to "Mix flour, butter, and tofu," the chef actually follows the new list because they aren't distracted by the name of the dish.

The Surprising Twist: Thinking Can Make It Worse

One of the most interesting findings was about "reasoning." The author asked the AI to "think step-by-step" before answering.

  • The Finding: In the first two tests, asking the AI to think actually made it more stubborn. The AI would write a long, logical paragraph explaining why it should follow the "Tit-for-Tat" strategy (a famous strategy for the original game), completely ignoring the fact that the rules had changed.
  • The Analogy: It's like a student who memorized the answer key for a math test. If you change the numbers in the problem but ask them to "show their work," they might write out a perfect, logical explanation of why the old answer is right, completely missing that the question itself has changed.

The Bottom Line

The paper concludes that AI models are like students who have studied too hard for a specific exam. They know the "standard" answers so well that if you slightly change the question, they might not notice.

  • The Problem: When we give AI tasks using familiar names or scenarios (like "Prisoner's Dilemma"), the AI relies on its "presumptions" about those scenarios rather than the specific instructions you just gave it.
  • The Solution: To get the AI to follow the actual rules you give it, you should describe the task in a neutral, abstract way. Don't tell it what the game is; just tell it what the rules are.

In short: If you want an AI to adapt to a new situation, don't remind it of the old one. Speak in "variables" (X and Y) rather than "famous names."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →