← Latest papers
💻 computer science

The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost

This study demonstrates that while strategic model selection and increased reasoning effort significantly enhance automated scoring accuracy for high school mathematics conversations, ensembling via self-consistency yields no substantial gains, with specific low-cost models offering the optimal balance between performance and expense.

Original authors: Scott Frohn

Published 2026-05-01
📖 4 min read☕ Coffee break read

Original authors: Scott Frohn

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher grading hundreds of short math conversations where students explain their thinking. You want to use an AI robot to do this grading for you, but you have two big worries: Will the robot be accurate? and Will it cost too much money?

This paper is like a lab experiment where researchers tested different "settings" on various AI robots to see which ones were the best at grading these math conversations. They compared robots from two big companies (OpenAI and Google) and tested two main strategies:

  1. The "Committee" Strategy (Self-Consistency): Asking the same robot to grade the same answer multiple times and taking a vote.
  2. The "Deep Thinker" Strategy (Reasoning Effort): Telling the robot to "think harder" before giving a grade.

Here is what they found, explained simply:

1. The "Committee" Strategy Didn't Help Much

The Idea: If one robot is a bit confused, maybe asking it to grade the same answer five times and taking the majority vote will smooth out the mistakes. It's like asking five friends to guess the answer to a riddle; if they all agree, you feel more confident.

The Reality: The researchers found that these AI robots are actually very stubborn. Once they decide on an answer, they stick to it, even if they are wrong.

  • The Analogy: Imagine a robot that is convinced the sky is green. If you ask it 10 times, it will say "green" 10 times. Asking it more times doesn't make it realize the sky is actually blue; it just makes you pay for 10 answers instead of one.
  • The Result: Asking the robot to vote on itself (from 1 time up to 7 times) did not improve accuracy. It just increased the cost. The only thing that helped slightly was making the robot a little more "random" (changing a setting called temperature), but simply asking it to vote more didn't fix the errors.

2. The "Deep Thinker" Strategy Was Hit or Miss

The Idea: Newer AI models can be told to "think step-by-step" before answering, similar to how a human might scratch out a math problem on paper before writing the final answer.

The Reality: This worked, but it depended entirely on which robot you used.

  • The "Over-Thinkers": For some cheaper, smaller robots, telling them to "think harder" actually made them worse at grading. They started over-analyzing simple answers and getting confused.
  • The "Smartest" Robot: The most powerful robot (Gemini 3.1 Pro Preview) was already so good that telling it to think harder didn't really change its score. It was accurate whether it thought for a second or a minute.
  • The General Trend: Overall, letting the models think a bit more did help accuracy, but the improvement wasn't a straight line. Sometimes thinking more helped, sometimes it didn't.

3. The "Price vs. Performance" Sweet Spot

The researchers drew a map (called an "efficiency frontier") to show the best balance between how good the grade was and how much it cost.

  • The Cheap Winners: The smallest, cheapest robots (GPT-5.4 Nano and Mini) with zero "deep thinking" turned out to be the best value. They were surprisingly accurate and cost pennies per grading task.
  • The Expensive Champion: The most powerful robot (Gemini 3.1 Pro Preview) gave the absolute highest accuracy, but it cost about 4 to 15 times more than the cheap robots.
  • The Middle Ground: Telling the cheaper robots to "think harder" usually just made them more expensive without making them much more accurate.

The Bottom Line

If you want to grade student math conversations automatically:

  • Don't ask the AI to vote on its own answers multiple times; it wastes money without fixing mistakes.
  • Do pick the right robot for your budget. If you need the absolute best accuracy and money is no object, use the most powerful model. But if you need to grade thousands of papers cheaply, a smaller, "lazy" (non-thinking) robot is often just as good and much cheaper.

Important Note: This study only looked at high school math conversations with simple "yes/no" checklists. The results might be different for essays, history questions, or more complex grading scales. Also, the prices mentioned are based on current API rates, which can change.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →