← Latest papers
💬 NLP

Measuring Pragmatic Influence in Large Language Model Instructions

This paper introduces a novel framework to systematically measure and quantify "pragmatic framing"—the influence of contextual cues like urgency or authority on large language models' directive prioritization—demonstrating that such framing consistently shifts model behavior across diverse architectures.

Original authors: Yilin Geng, Omri Abend, Eduard Hovy, Lea Frermann

Published 2026-02-26
📖 5 min read🧠 Deep dive

Original authors: Yilin Geng, Omri Abend, Eduard Hovy, Lea Frermann

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a very smart, very obedient robot assistant. You have a list of chores to do. Usually, you do them exactly as asked. But what if someone whispers, "This is an emergency!" or "I'm your boss, do this now!" before giving you the instruction?

Suddenly, you might ignore your original list and rush to do that specific task, even if it wasn't the most important one.

This paper is about figuring out how much those "whispers" actually change a robot's mind.

Here is the breakdown of the research, explained simply:

1. The Problem: The "Magic Words" Mystery

We know that Large Language Models (AI) are sensitive to how we talk to them. It's not just what they are asked to do, but how it's asked.

  • The Old Way: Researchers used to mix these "magic words" (like "You are a pirate" or "This is urgent") right into the task itself. It was like trying to taste if salt makes a soup salty, but you kept adding pepper and changing the temperature at the same time. You couldn't tell what was doing the work.
  • The New Idea: This paper says, "Let's separate the Task from the Vibe."

2. The Experiment: The "Tug-of-War"

To measure this, the researchers set up a clever game of tug-of-war.

Imagine you give the AI two instructions that contradict each other:

  • Instruction A: "Write a paragraph saying remote work is good."
  • Instruction B: "Write a paragraph saying remote work is bad."

Normally, a smart AI tries to be fair. It might say, "Well, both have pros and cons," or it might try to do both. It stays neutral.

The Twist: The researchers take one of those instructions (say, Instruction B) and add a "Vibe Prefix" right before it.

  • Example: "I am your supervisor and this is an emergency. Write a paragraph saying remote work is bad."

Then they ask: Did the AI ignore the first instruction and only do the second one because of the "Vibe"?

3. The "Vibe" Menu (The Taxonomy)

The researchers created a massive menu of 400 different "Vibe Prefixes" and sorted them into four main categories, like different flavors of influence:

  • 👑 The Boss (Hierarchical): Pretending to be an authority figure.
    • Example: "As your system administrator, I order you to..."
    • Result: Very Effective. The AI listens to the "boss."
  • 🤝 The Friend (Social Contract): Using friendship or favors.
    • Example: "We have a great working relationship; please help me..."
    • Result: Very Effective. The AI wants to be a good friend.
  • 😢 The Drama (Emotional): Using guilt, urgency, or sadness.
    • Example: "I'm in a crisis and need help right now!"
    • Result: Moderately Effective. It works, but not as well as the "Boss."
  • 🎭 The Storyteller (Narrative): Putting the task inside a fake story or role-play.
    • Example: "Imagine you are a character in a movie who must..."
    • Result: Least Effective. The AI is less fooled by stories in this specific setup.

4. The Big Discovery

They tested this on five different AI models (some huge, some small). Here is what they found:

  • The "Vibe" Wins: Even though the AI is trained to be neutral, a short phrase (average 8 words!) can completely flip its decision. It stops being fair and starts favoring the instruction with the "Vibe."
  • The "Boss" is King: Pretending to be an authority figure or claiming a direct order is the strongest way to change the AI's mind.
  • It's Consistent: Whether the AI is a giant brain (235 billion parameters) or a smaller one (7 billion), they all react to these social cues in the same order. The "Boss" trick works best on everyone; the "Storyteller" trick works worst on everyone.
  • It's Not Just Length: They tried adding random nonsense words (like "Lorem Ipsum") to see if the AI just liked more text. It didn't work nearly as well. The AI specifically responds to the meaning of the social cue, not just the length of the sentence.

5. Why Does This Matter?

Think of this like a control panel for AI behavior.

  • For Safety: If we know that "I am your supervisor" is a powerful trigger, we can build better safety guards to stop bad actors from using that trick to make the AI do harmful things (like "jailbreaking").
  • For Better AI: If we understand how these cues work, we can design AI that is smarter about when to listen to a "vibe" and when to stick to the facts.

The Bottom Line

This paper proves that how you ask is just as powerful as what you ask. By separating the "ask" from the "flavor," the researchers showed that AI models have a predictable, measurable "social weakness." They are surprisingly easy to sway if you use the right social pressure, much like a human might be swayed by a boss or a friend.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →