← Latest papers
🤖 AI

Chat-Based Support Alone May Not Be Enough: Comparing Conversational and Embedded LLM Feedback for Mathematical Proof Learning

In a study of 148 undergraduate students using the GPTutor system, findings indicate that while structured, embedded proof-review feedback did not harm learning outcomes, reliance on chatbot-based support—particularly answer-seeking behavior—was negatively associated with subsequent exam performance, suggesting that conversational AI alone may not effectively foster transfer to independent mathematical proof assessments.

Original authors: Eason Chen, Sophia Judicke, Kayla Beigh, Xinyi Tang, Isabel Wang, Nina Yuan, Zimo Xiao, Chuangji Li, Shizhuo Li, Reed Luttmer, Shreya Singh, Maria Yampolsky, Naman Parikh, Yvonne Zhao, Meiyi Chen, Sca
Published 2026-04-02
📖 4 min read☕ Coffee break read

Original authors: Eason Chen, Sophia Judicke, Kayla Beigh, Xinyi Tang, Isabel Wang, Nina Yuan, Zimo Xiao, Chuangji Li, Shizhuo Li, Reed Luttmer, Shreya Singh, Maria Yampolsky, Naman Parikh, Yvonne Zhao, Meiyi Chen, Scarlett Huang, Anishka Mohanty, Gregory Johnson, John Mackey, Jionghao Lin, Ken Koedinger

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are learning how to write a complex, logical story (a mathematical proof). You have a new, super-smart AI assistant available to help you. But this assistant comes in two very different "personas":

  1. The Chatty Friend (The Chatbot): You can text back and forth. You ask, "How do I start?" and it might give you hints, or if you push hard enough, it might just tell you the answer.
  2. The Strict Editor (The Proof-Review Tool): You must write your story first. Then, you paste it in. The tool doesn't talk to you; it just highlights specific sentences in red and writes comments in the margins saying, "This logic doesn't follow," or "Great job here." It refuses to write the story for you.

The researchers at Carnegie Mellon University wanted to see which of these two helpers actually made students better at math. They gave 148 students access to a system called GPTutor that had both tools.

Here is what they found, broken down simply:

1. The "Homework vs. Exam" Surprise

When students had access to the AI, their homework scores went up. It was like having a safety net; they could draft their proofs, get feedback, fix mistakes, and turn in better work.

However, when it came time for the midterm exams (where they couldn't use the AI), the students who used the tool did not score higher than those who didn't.

  • The Analogy: It's like practicing basketball with a coach who holds the hoop for you. You make a lot of shots during practice (homework), but when you play a real game without the coach (the exam), you miss just as often as everyone else. The practice didn't translate to real skill.

2. The Danger of the "Chatty Friend"

The study dug deeper to see how students used the tools. They found a scary pattern with the Chatbot:

  • Students who felt less confident in their math skills (low "self-efficacy") used the chatbot the most.
  • Crucially, the more these students chatted with the bot, the worse they did on the exams.
  • The Metaphor: The chatbot was like a "crutch." Students with shaky legs leaned on it so much that their own muscles (their own reasoning skills) never got strong enough to walk on their own. Even though the bot was programmed to give hints, students often bypassed the hints and just asked for the answer, short-circuiting their own learning.

3. The Safety of the "Strict Editor"

The Proof-Review Tool told a different story.

  • Students who used this tool also tended to be the ones who were struggling.
  • But, using this tool did not hurt their exam scores. It didn't help them skyrocket, but it didn't drag them down either.
  • The Metaphor: This tool was like a "mirror." You had to stand in front of it and look at your own work. It pointed out your flaws, but you still had to fix them yourself. Because you were forced to do the heavy lifting, you didn't lose your ability to think independently.

The Big Takeaway

The paper's main message is: Just having a smart AI chatbot isn't enough to teach you hard things.

If you let an AI chat with you freely, you might get lazy and stop thinking for yourself. You might get good at using the tool, but you won't get good at the subject.

However, if you force the AI to act like a strict editor—where you have to do the work first, and the AI only critiques your specific mistakes—you avoid the trap of over-reliance. You still struggle (which is good for learning), but you don't outsource your brain to the machine.

In short: Don't let the AI write the story for you. Let it edit your story, but make sure you are the one holding the pen.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →