Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses
This paper proposes and evaluates a new framework for assessing large language models' ability to generate responses with varying language complexity, revealing that current models struggle to consistently adjust their output style to match user needs, with the best-performing model succeeding in only 46% of cases.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are at a restaurant where the chef (the AI) is supposed to cook the same dish (an answer to a question) but serve it to five different people: a 5-year-old, a college student, a PhD student, a postdoc, and a senior professor.
The paper asks a simple but tricky question: Can the chef actually change the recipe enough so that each person gets a meal that truly matches their taste, or does the chef just serve the same plate with slightly different garnishes?
Here is the breakdown of the research, explained simply:
1. The Problem: The "One-Size-Fits-All" Menu
Right now, when we talk to AI chatbots, we usually get one answer. If you want a simpler answer, you have to type a new prompt like, "Explain this like I'm five." If you want it more complex, you type again.
The researchers wanted to see if we could build a better interface where you have a slider (like a volume knob). You could slide it from "Simple" to "Expert," and the AI would instantly rewrite the answer for you.
But before we build these sliders, we need to know: Is the AI actually good at turning that knob? Does it know how to make the text truly different, or does it just shuffle the words around?
2. The Test: The "Taste Test"
The researchers did two things to find out:
Step A: The Human Taste Test (The Formative Study)
They asked 16 people (mostly scientists and students) to try a prototype slider. They found that people loved having control. They wanted to be able to dial down the "jargon" (fancy words) or the amount of information.- The Catch: The participants noticed that sometimes the "Simple" version and the "Medium" version felt almost identical. The slider wasn't moving the needle enough.
Step B: The Robot Taste Test (The Model Evaluation)
They took 98 difficult science questions and asked 5 different AI models (like GPT-5, Claude, and DeepSeek) to generate 5 versions of the answer for each question, ranging from "College Student" to "Senior Researcher."
3. The Ingredients of Complexity
To measure if the AI actually changed the "flavor," they looked at three specific ingredients:
- Jargon: Are there fewer fancy, technical words?
- Information: Is there less dense technical detail?
- Length: Is the text shorter? (Usually, simpler answers are shorter, but not always).
4. The Results: The Chef is Confused
Here is what they found, using a simple analogy:
The Length Knob Works: If you asked the AI to make the text "simpler," it almost always made the text shorter. It's like the chef always serving a smaller portion when asked.
The Jargon and Info Knobs are Broken: This is the big problem. When the researchers asked the AI to make the text more complex (for the experts), the AI often accidentally made it simpler or just kept it the same.
- The Stat: Even the best AI model (Claude Sonnet 4.5) only got the "direction" right about 46% of the time. That's barely better than flipping a coin!
- The "Elaborative Simplification" Trap: Sometimes, when the AI tried to make a complex answer "simpler," it didn't remove the hard words. Instead, it just added more words to explain the hard words in a long, rambling way. It made the text longer, but not actually easier to understand.
The "First Step" Problem: The AI was pretty good at changing the answer when going from "Level 1" (College Student) to "Level 2" (Junior PhD). But as they tried to go from "Level 3" to "Level 4" to "Level 5," the AI got confused and the differences between the answers became random.
5. The Conclusion
The paper concludes that while AI is getting very good at writing text, it is currently bad at being a "shape-shifter."
If you build a slider interface today, the user might slide it from "Simple" to "Complex," and the AI might just give them the exact same answer, or an answer that is actually less complex than before.
The main takeaway: We cannot just assume AI can handle interactive controls (like sliders) yet. We need to teach these models how to reliably change their "voice" and "depth" before we give users the controls to do it themselves.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.