← Latest papers
🔢 mathematics

Auditing Generative AI Error Patterns in Applied Differential Calculus: Evidence from ChatGPT and Gemini

This descriptive-evaluative study audits ChatGPT and Gemini's performance in applied differential calculus, revealing that while both models can execute isolated procedures, they frequently fail in translating narrative conditions to functions, handling boundary constraints, and maintaining arithmetic consistency, thereby suggesting their solutions should be used as critique artifacts rather than unquestioned answers in education.

Original authors: Polemer M. Cuarto

Published 2026-07-10
📖 6 min read🧠 Deep dive

Original authors: Polemer M. Cuarto

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've hired two super-smart, incredibly chatty robots, ChatGPT and Gemini, to help you solve tricky math puzzles about real-world problems. You ask them to figure out how to make the most money for a business, how fast a particle is moving, or how high a rocket will fly. You expect them to be perfect calculators, right?

Well, a new study by researcher Polemer M. Cuarto decided to put these robots to the test. The findings are a bit like discovering that your favorite robot chef can chop vegetables perfectly but keeps forgetting to add the salt, or worse, thinks the recipe calls for sugar instead of salt.

The Main Discovery: The "Fluent but Flawed" Trap

The study found that while these AI tools are great at doing the mechanical math steps—like taking a derivative or an integral—they often stumble when they have to translate a story into a math equation or remember the starting conditions.

Think of it like this: If you ask a robot to build a house, it might lay the bricks (the math operations) perfectly. But if it misunderstands your blueprint and thinks you wanted a castle instead of a cottage (the math model), the whole house is wrong, no matter how neatly the bricks are stacked.

The paper suggests that these tools are not reliable "answer keys" you can just copy from. Instead, they are better used as practice partners for finding mistakes. The researchers argue that the most dangerous thing about AI isn't that it gets the math wrong; it's that it gets the math wrong while sounding completely confident and professional.

The Three Big Stumbles

The study looked at three specific types of problems, and the robots tripped over the same hurdles in each:

1. The Business Mix-Up (Semantic-to-Algebraic Translation)
Imagine a store owner says, "If I lower the price by $10, I sell 20 more items."

  • The Expert Way: This means the price goes down as sales go up. The math slope is negative.
  • The AI Way: Both ChatGPT and Gemini sometimes got this backward. They built a model where lowering the price somehow decreased sales, or they flipped the relationship entirely.
  • The Result: They then did perfect math on a broken model. It's like calculating the fastest route to a destination that doesn't exist. The paper notes that in the business problem, the AI produced a "positive-slope demand relation," which is economically nonsense (implying higher prices lead to more sales in a way that contradicts the problem).

2. The "Ghost" Starting Point (Boundary-Condition Tracking)
In physics, if you drop a ball from a 140-foot tower, you have to remember that 140-foot height.

  • The Expert Way: You keep that "140" in your equation the whole time.
  • The AI Way: ChatGPT often just forgot the starting height or set it to zero. It's like the robot looked at the ball, saw it falling, and decided, "Oh, it must have started from the ground!" even though you told it it started from a tower.
  • The Result: The final answer for when the ball hits the ground was wrong because the robot was solving a different problem than the one you asked.

3. The Arithmetic Slip-Up (Verification and Context)
Sometimes the setup was okay, but the final number crunching went wrong.

  • The AI Way: Gemini was good at setting up the rocket equation but stumbled when solving the final quadratic formula (the math to find the time). It also sometimes added weird, irrelevant explanations about gravity that weren't needed.
  • The Result: The answer looked plausible but was numerically shaky. The paper highlights that these tools can produce "fluent but flawed explanations," inserting irrelevant info that distracts from the actual error.

What the Paper Says You Should Not Do

The researchers are very clear about what not to do. They argue against using these AI tools as an "unquestioned answer source."

  • Don't just copy the solution because it looks neat and has all the right formulas.
  • Don't assume that because the AI sounds confident, it is right.
  • Don't think that if the final number looks close, the whole process is valid.

The paper explicitly rules out the idea that these tools are ready to replace human judgment in calculus. In fact, the study suggests that relying on them without checking is dangerous because the errors are hidden inside a "polished" presentation.

How Sure Are We?

The authors are careful not to call this a "final verdict" on all AI. They admit their study is based on a small set of specific problems (business, particle motion, and vertical motion) solved in early 2026. They suggest that these error patterns are real and observable in these specific tasks, but they don't claim to have measured every possible math problem the AI could ever face.

They found that the tools showed "partial procedural competence" (they can do the steps) but "weaker reliability" in modeling and checking. The conclusion is a strong suggestion that we need to change how we teach math: instead of asking students to just "get the answer," we should teach them to be math auditors.

The New Game Plan: Be the Detective

So, what's the takeaway for a curious student?

Don't treat the AI like a magic oracle. Treat it like a student who is really good at writing but bad at thinking.

  • The Strategy: Ask the AI to solve a problem, then play detective.
  • The Mission: Find the first step where the robot went wrong. Did it misunderstand the story? Did it forget the starting height? Did it mess up the arithmetic?
  • The Goal: Fix the mistake and explain why the AI was wrong.

The paper suggests that this "audit" approach turns the AI from a cheat sheet into a powerful learning tool. By hunting for the robot's mistakes, you actually learn the math better than if you had just solved it yourself.

In short: The robots are great at doing the heavy lifting of calculation, but they are terrible at understanding the story. You have to be the one holding the map, checking the compass, and making sure they don't accidentally build a castle when you asked for a cottage.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →