Using LLMs to Detect Growth in Computational Thinking in Introductory Physics
This study demonstrates that Large Language Models can effectively scale the assessment of student growth in computational thinking within introductory physics courses by mirroring human evaluations of written explanations for data practices and problem-solving, though both methods face challenges in reliably evaluating complex constructs like systems thinking.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of physics not just as a collection of formulas on a chalkboard, but as a massive, interactive video game where the laws of nature are the code. For decades, teachers have tried to get students to stop just memorizing the rules and start writing the code themselves. This skill is called "Computational Thinking." It's not just about typing fast or knowing computer syntax; it's about breaking a giant, messy real-world problem (like a car crash or a planet's orbit) into tiny, logical steps a computer can understand. It's like being a translator who speaks both "Human" and "Machine," figuring out how to describe a falling apple so a robot can predict exactly where it will land.
The big question for educators is: How do we know if students are actually getting better at this translation job? Traditionally, teachers have used multiple-choice tests, which are easy to grade but terrible at showing how a student thinks. To see the real thinking, students have to write out their answers in words. But here's the catch: reading hundreds of handwritten essays is exhausting, slow, and expensive. It's like trying to grade a million essays by hand while the clock is ticking. Now, a new kind of "super-reader" has arrived: Artificial Intelligence (AI), specifically Large Language Models (LLMs). These are the same types of smart computers that can write stories or chat with you. The burning question is: Can these AI bots read student essays and spot growth in their thinking skills as well as a human teacher can, but without the coffee breaks?
This paper, written by researchers at Purdue University, sets out to answer that question. They wanted to see if an AI could act as a giant, tireless teaching assistant to track how students' computational thinking grows over a semester. They didn't just ask the AI to grade; they asked it to look for specific "superpowers" in the students' writing: Can they handle data? Can they break problems down? Can they understand how complex systems work together?
To test this, the researchers gathered 936 students in a physics class where they learned to code in Python. Before the class started and after it ended, the students answered open-ended questions about physics problems, like interpreting a graph of a bouncing spring or designing a simulation for a sled sliding down a hill. First, a team of human experts read a small sample of these essays and gave them scores based on a strict rubric. This created a "gold standard" of what a human thinks is a good answer. Then, they fed the exact same essays to an AI model (specifically a multimodal version of GPT-5.4-mini) and asked it to grade them using the same rules.
The results were surprisingly promising. The AI turned out to be a very capable assistant. For clear-cut skills like handling data (Data Practices) and breaking down code steps (Computational Problem-Solving), the AI's grades matched the human experts' grades almost perfectly. When the researchers looked at the whole group of 936 students, the AI successfully spotted the same big trends the humans found: students got significantly better at handling data and solving computational problems by the end of the semester. In fact, the AI confirmed that students made huge leaps in these areas, with the data skills showing a massive improvement.
However, the story isn't a perfect "AI wins everything" tale. The researchers found that both humans and the AI struggled with the most complex skill: "Systems Thinking." This is the ability to see how all the tiny parts of a system interact to create a big picture. Because this concept is so fuzzy and hard to pin down, even the human experts disagreed with each other sometimes, and the AI struggled to match them too. The paper suggests that the AI isn't "dumb" here; rather, the task itself is just really hard to define clearly.
There was also a funny twist with one specific question about designing a sled simulation. The students were already so good at this topic before the class even started that they couldn't get much better. It was like asking a professional chef to improve their knife skills when they were already using the sharpest knives in the world. The scores hit a "ceiling," meaning there was no room for growth to be measured, and the AI correctly identified that no progress happened there.
In the end, the study shows that AI is a viable, powerful tool for checking how students learn to think like computational physicists. It can scale up to read thousands of essays and find the same patterns humans do, saving teachers from drowning in paperwork. But it also serves as a reminder that AI is only as good as the rules we give it. For the messy, complex parts of thinking where even humans disagree, we still need careful design and human oversight. The AI isn't replacing the teacher; it's handing them a super-powered magnifying glass to see the growth that was already happening.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.