← Latest papers
🤖 AI

Temporal Stability and Few-Shot Prompting in Math Task Assessment

This longitudinal study demonstrates that while model version updates alone yield inconsistent results in classifying the cognitive demand of math tasks, few-shot prompting significantly and reliably improves the performance of both general-purpose and education-specific AI tools, suggesting that prompt engineering is a more effective strategy for enhancing AI in educational contexts than relying solely on passive model improvements.

Original authors: Danielle S. Fox, Brenda L. Robles, Elizabeth DiPietro Brovey, Christian D. Schunn

Published 2026-05-29
📖 5 min read🧠 Deep dive

Original authors: Danielle S. Fox, Brenda L. Robles, Elizabeth DiPietro Brovey, Christian D. Schunn

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you hire two different "AI teaching assistants" to help you sort a pile of math homework problems. Your goal is to separate them into two piles: "Easy, routine drills" and "Hard, thinking-heavy challenges." You give them a rulebook (called the Task Analysis Guide) to help them decide.

This study is like a long-term check-up on these two assistants to see if they stay consistent over time and if giving them a "cheat sheet" of examples helps them do a better job.

Here is what happened, explained simply:

The Two Assistants

The researchers tested two different AI tools:

  1. Gemini: A "general-purpose" assistant. Think of this like a very smart, well-read librarian who knows a little bit about everything, including math.
  2. Coteach: An "education-specific" assistant. Think of this like a veteran math teacher who has spent years grading papers and knows the specific quirks of school math.

Part 1: The "Time Travel" Test (Temporal Stability)

The researchers asked both assistants to sort the math problems on Day 1. Then, they waited seven weeks and asked them to sort the exact same problems again.

In the real world, software updates happen constantly. It's like if your favorite video game got a patch update; sometimes the game gets better, but sometimes a character you liked suddenly moves differently or gets weaker.

  • What happened to the General Assistant (Gemini)?
    Its overall score stayed exactly the same (58% correct). However, if you looked closely, it had changed its mind on specific problems. It got three new problems right that it missed before, but it got three old problems wrong that it had previously solved.

    • The Analogy: Imagine a judge who gives you the same sentence on average, but swaps which specific crimes get the death penalty and which get life in prison. The "average" looks the same, but the specific outcome for your case has changed unpredictably. This is called "prompt drift."
  • What happened to the Teacher Assistant (Coteach)?
    This one got worse. Its score dropped significantly from 75% correct down to 50%.

    • The Analogy: Imagine a veteran teacher who suddenly starts forgetting how to grade basic tests. They didn't just swap their answers; they actually lost some of their ability to do the job correctly.

The Big Takeaway: Just because an AI tool gets a "new version" or updates in the background doesn't mean it gets better at your specific job. Sometimes it stays the same but acts differently, and sometimes it actually gets worse.

Part 2: The "Cheat Sheet" Test (Few-Shot Prompting)

After the seven-week wait, the researchers tried a new trick. Instead of just asking the assistants to sort the problems, they gave them a "cheat sheet." This sheet showed them two perfect examples of an "Easy" problem and two perfect examples of a "Hard" problem, along with the correct labels for each.

This is called Few-Shot Prompting. It's like saying to a new employee, "Here is a sample of a good report and a sample of a bad report. Now, look at this new pile and sort them like these."

  • The Results:
    • Gemini (The Librarian): Got better! Its score went up from 58% to 67%.
    • Coteach (The Teacher): Got much better! It bounced back from its low 50% score all the way up to 75%, returning to its original "expert" level.

The Big Takeaway: Giving the AI a few clear examples (the cheat sheet) helped it perform better than waiting for the software company to release a new version. In fact, the "cheat sheet" fixed the teacher assistant's mistakes completely.

The "Hard Stuff" Problem

There was one catch. The AI tools were great at sorting the "Easy" routine problems. But they really struggled with the "Hard, thinking-heavy" problems (called "Doing Mathematics").

Even with the cheat sheet, the AI kept misclassifying the hardest problems.

  • The Analogy: It's like a student who is great at memorizing multiplication tables but gets confused when asked to solve a word problem that requires them to invent a new way to think. The AI tends to look at the surface words (like "calculate" or "show work") rather than understanding the deep thinking required. It tries to find a shortcut instead of doing the hard mental work.

Summary of the Paper's Claims

  1. AI is unstable: Even over a short time (7 weeks), AI tools can change how they answer questions, sometimes getting worse, sometimes just acting unpredictably.
  2. Updates aren't magic: Newer versions of AI don't automatically mean better performance on specific school tasks.
  3. Examples help: Giving the AI a few clear examples (Few-Shot Prompting) is a powerful way to fix mistakes and improve accuracy, often more effectively than waiting for software updates.
  4. Specialized tools need help: The education-specific tool (Coteach) was more sensitive to changes (it got worse faster) but also responded better to the "cheat sheet" than the general tool.
  5. Hard thinking is still hard: AI still struggles to identify the most complex math tasks, often confusing them with simpler ones, even when given examples.

The paper concludes that if you are a teacher or researcher using AI, you shouldn't just rely on the tool being "new." You need to actively check if it's still working right and be ready to give it clear examples to guide its thinking.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →