Using an LLM to Investigate Students' Explanations on Conceptual Physics Questions
This study demonstrates that large language models can effectively and accurately scale the assessment of students' written physics explanations, revealing deeper misconceptions that traditional multiple-choice formats fail to capture.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a teacher grading a massive physics class with over 1,000 students. You want to know not just if they got the right answer, but why they think that way. Traditionally, you'd have to give them a multiple-choice test because it's the only way to grade 1,000 papers quickly. But multiple-choice tests are like a locked door: they tell you if a student can pick the right key, but they don't show you how the student is trying to unlock the door. They might pick the right key by guessing, or they might be using a completely wrong method that just happens to lead to the right door.
This paper is about trying to open that door using a new kind of "digital assistant" called a Large Language Model (LLM), specifically a version of AI called GPT-4o.
The Experiment: Swapping the Door for a Window
The researchers took a standard physics test called the "Energy and Momentum Conceptual Survey" (EMCS). Usually, this test is multiple-choice. For this study, they took three specific questions that students historically struggle with and turned them into essay questions.
Instead of circling "A, B, C, or D," students had to write out their answers using a specific recipe:
- Claim: What is your answer?
- Evidence: What facts do you have?
- Reasoning: How do those facts prove your answer?
This is like asking a student not just to say "The car is fast," but to explain, "The car is fast because the engine is powerful (evidence) and power creates speed (reasoning)."
The Test: Human vs. Machine
The researchers wanted to see if the AI could grade these essays as well as a human teacher.
- The Human Grader: A real physics expert read a random sample of 231 essays.
- The AI Grader: The LLM read all 1,131 essays.
The Result: The AI and the human agreed almost perfectly. They only disagreed on whether a student was "correct" or "incorrect" about 0% to 3% of the time. It's as if you asked two different chefs to taste a soup and rate its saltiness; they would give you nearly the same score. The researchers concluded that the AI is a reliable "digital grader" that can handle the workload of a huge class without getting tired.
The Big Discovery: Finding Hidden Clues
Here is where the paper gets really interesting. The researchers asked the AI to do something a multiple-choice test can't do: categorize the mistakes.
When a student gets a multiple-choice question wrong, they just pick the wrong option (a "distractor"). The test designers assume that wrong option represents a specific misunderstanding. But the AI, reading the essays, found that students were making mistakes the test designers never thought of.
The Analogy:
Imagine a detective looking for a thief.
- The Multiple-Choice Test is like a lineup of three suspects. The detective picks one, but maybe the real thief wasn't even in the lineup.
- The AI Analysis is like the detective reading the thief's diary. The diary reveals the thief was thinking about a completely different crime, or using a tool the detective didn't know existed.
Real Examples from the Paper:
- Question 5: The multiple-choice options suggested students were confused about "momentum" or "Newton's Laws." However, the AI found that many students were actually confused about the type of collision (elastic vs. inelastic) in a way the test never asked about. They were mixing up concepts like "bouncing" and "sticking" in ways the multiple-choice options couldn't capture.
- Question 16 & 23: The AI found students who got the right final answer but used completely wrong logic to get there. In a multiple-choice test, these students would get a "correct" score. The AI, however, spotted that their reasoning was flawed (like using the wrong physics principle) and flagged it as a "misapplication."
The Bottom Line
The paper claims two main things:
- Reliability: An AI can grade physics essays just as accurately as a human teacher, making it possible to grade written explanations for huge classes without burning out the staff.
- Depth: The AI can uncover "hidden" misunderstandings that multiple-choice tests miss. It acts like a microscope, showing educators the specific, messy ways students are thinking about physics, rather than just a simple "Right/Wrong" score.
The researchers are careful to say they only tested this on three questions and a specific group of students. They aren't claiming this fixes all education problems yet, but they have proven that the "digital assistant" can see things the "multiple-choice door" keeps hidden.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.