← Latest papers
💬 NLP

Investigating LLM's Problem Solving Capability -- a Study on Statics Questions

This study evaluates Large Language Models' capabilities in solving statics problems through a model distillation approach, revealing that while they perform well on text-only questions, their accuracy significantly declines when diagrams and multi-step reasoning are required due to challenges in consistently applying visual information rather than image recognition limitations.

Original authors: Tanner Culleton, Hung-Fu Chang

Published 2026-06-26
📖 4 min read☕ Coffee break read

Original authors: Tanner Culleton, Hung-Fu Chang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, well-read student named "ChatGPT" who has read almost every book in the library. You want to test if this student can actually solve tough engineering homework problems, specifically in a subject called Statics (which deals with how forces balance on stationary objects, like bridges or cranes).

Here is what the researchers did, explained simply:

1. The "Copy-Paste" Problem

First, the researchers tried the obvious thing: they took real textbook questions and asked ChatGPT to solve them.

  • The Result: It was a disaster. The student got confused, especially when the problems involved 3D drawings. It would mix up directions (like saying "up" when it meant "right") and get the math wrong.
  • The Analogy: It was like asking a student to read a map written in a language they barely know. They knew the words, but they couldn't figure out the directions.

2. The "Reverse Engineering" Trick (Model Distillation)

Since the textbook questions were too hard, the researchers tried a clever trick called Model Distillation. Instead of asking the student to solve other people's questions, they asked the student to write its own questions.

  • The Logic: They figured, "If ChatGPT can write a good question, it probably knows the answer to that specific type of question because it 'learned' it from its own training data."
  • The Process: They asked ChatGPT to generate 50 statics problems. They noticed the student kept writing the same simple, 2D (flat) problems over and over. So, they picked the best 25 unique questions that ChatGPT seemed to understand well.

3. The Three Levels of the Test

The researchers created three different "exams" using those 25 questions to see how the student's brain worked:

  • Exam A (Text-Only): They asked ChatGPT to solve the questions it wrote, using only words.
    • Result: Perfect Score (100%). The student knew the answers because it wrote the questions!
  • Exam B (With Pictures): They took those same word-only questions and drew simple diagrams for them.
    • Result: Score dropped to 68%. The student could still read the words, but the picture threw it off. It struggled to connect the drawing to the math.
  • Exam C (The "New Numbers" Test): They took the diagrams from Exam B and changed all the numbers (e.g., changing a force from 10 lbs to 15 lbs).
    • Result: Score dropped to 60%. This proved the student wasn't just memorizing the answers. It was trying to solve the problem fresh, but it kept making mistakes in the middle of the process.

4. What Did They Learn?

The researchers compared ChatGPT to another AI student named "Gemini."

  • The Big Discovery: The problem wasn't that the AI couldn't "see" the pictures. The problem was that it couldn't think in steps.
  • The Analogy: Imagine a chef who can read a recipe perfectly (Text-Only). If you hand them the ingredients and a picture of the dish (Diagram), they can still cook it. But if you ask them to cook a complex, multi-course meal where they have to adjust the recipe halfway through (Multi-step reasoning), they start dropping ingredients or forgetting the next step.
  • The Conclusion: The AI is great at simple, one-step math. But when it has to look at a picture, extract information, and then use that information in three or four different steps to get a final answer, it gets lost. It follows the right method but often gets the wrong number at the end.

Summary

The paper concludes that while AI is getting very good at reading and writing, it is still a bit clumsy when it comes to visual reasoning and long, complex chains of logic. It's like a student who knows the theory perfectly but freezes up when asked to apply it to a real-world drawing with changing numbers.

Note: The researchers only tested these specific engineering problems. They did not claim this applies to medical diagnoses, legal advice, or other fields, nor did they suggest how to fix the AI for the future. They simply reported how the AI performed on this specific set of engineering puzzles.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →