← Latest papers
💬 NLP

Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked?

This study reveals that the format of questions, including the number of options and specific wording, significantly impacts Large Language Model performance on reasoning tasks, often causing a disconnect between the accuracy of reasoning steps and the correctness of the final answer.

Original authors: Seok Hwan Song, Mohna Chakraborty, Qi Li, Wallapak Tavanapong

Published 2026-04-29
📖 4 min read☕ Coffee break read

Original authors: Seok Hwan Song, Mohna Chakraborty, Qi Li, Wallapak Tavanapong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher trying to test how smart a new student (an AI) is at solving logic puzzles. You want to see if the student truly understands the math or logic, or if they are just guessing the right answer.

This paper is like a report card from a study where the researchers asked the same logic problems to five different AI students, but they asked the questions in three very different ways:

  1. The "Essay" Style (Short Answer): "Here is a problem. Solve it and tell me the answer."
  2. The "Multiple Choice" Style: "Here is a problem. Pick the right answer from these 5 or 11 options."
  3. The "True or False" Style: "Here is a problem and a specific answer. Is this answer True or False?"

The researchers wanted to know: Does the way you ask the question change how well the AI performs?

Here are the main takeaways, explained simply:

1. The "Guessing Game" Effect

The biggest surprise was that getting the right final answer doesn't always mean the AI did the right work.

  • The Analogy: Imagine a student taking a multiple-choice test. They might not know how to do the math, but they guess "C" and get it right. If you only look at the bubble sheet, they got a 100%. But if you look at their scratch paper, they wrote nonsense.
  • The Finding: The AI often "guessed" the right letter in a multiple-choice question even when its reasoning steps were wrong. Conversely, sometimes the AI did the perfect math but picked the wrong letter at the very end because it got confused by the options.
  • The Lesson: If you only check the final answer, you might think the AI is smarter than it actually is.

2. The "Menu" Problem (Multiple Choice)

When the researchers gave the AI a list of options to choose from, the number of options and the words used mattered a lot.

  • Too Many Choices: When the list grew from 5 options to 11 options, the AI's performance often dropped. It's like being offered 5 flavors of ice cream vs. 50; the AI got overwhelmed and started guessing more.
  • The "Something Else" Option: The researchers added a weird option called "Something else" (or "None of the above").
    • If "Something else" was the correct answer, the AI often struggled to pick it, even if it did the math right. It seemed to prefer picking a specific number, even if that number was wrong.
    • The AI also seemed to have a "favorite seat" bias. It would pick the first option or the last option more often, regardless of whether it was right.

3. The "Yes/No" Trap (True or False)

When asking "True or False" questions, the specific words used changed the results.

  • True vs. False: The AI was generally better at saying "True" when the answer was correct than saying "False" when the answer was wrong. It seemed to have a bias toward agreeing.
  • The Wording: Changing the prompt from "True or False" to "Yes or No" made the AI significantly worse at some tasks. It's like how a person might answer "Yes" to "Do you want to go?" but hesitate on "Is it true that you want to go?" The AI is sensitive to these tiny linguistic shifts.

4. The "Instruction" Confusion

The researchers tried to help the AI by adding instructions like, "Solve the problem first, then decide if it's True or False."

  • The Result: Sometimes this helped, but often it confused the AI or made it perform worse. The AI sometimes got so focused on the instruction that it forgot the actual logic of the problem.

Summary: What Should We Learn?

The paper concludes that we cannot just look at the final score to judge an AI's intelligence.

  • If you ask an AI a multiple-choice question, it might be "cheating" by guessing the right letter without understanding the logic.
  • If you ask it a True/False question, the specific words you use can trick it into being less accurate.
  • To truly know if an AI is smart, you have to check its "scratch paper" (its reasoning steps), not just the final bubble it filled in.

The researchers built a new set of tests (a benchmark) to help future developers understand these quirks, ensuring that when we test AI, we aren't just testing how good it is at guessing, but how good it is at thinking.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →