← Latest papers
🔬 physics

AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics

This study reveals that AI-based scoring systems systematically underestimate the conceptual understanding of linguistically weak physics students, a language bias that mirrors human expert tendencies and poses significant risks for multilingual learners in high-stakes assessments.

Original authors: Markus S. Feser, Paul L. Tschisgale

Published 2026-07-31
📖 5 min read🧠 Deep dive

Original authors: Markus S. Feser, Paul L. Tschisgale

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: "Does this student actually understand how the universe works?" The only clue you have is a handwritten note. This is the daily reality for physics teachers. They can't peek inside a student's brain to see the gears turning; they can only judge understanding through the words the student writes. But here's the tricky part: writing is a skill all its own. A student might have a brilliant, rock-solid understanding of why sound travels through a wall but fail to explain it because their sentences are clunky or their vocabulary is limited. Conversely, a student might write a beautifully polished essay that sounds smart but is actually full of nonsense. For decades, experts have known that human teachers sometimes get tripped up by this, accidentally giving lower scores to students who just aren't great writers, even if those students know their physics.

Now, enter the new detectives: Artificial Intelligence. Schools are starting to use AI to grade these essays because it's fast, consistent, and doesn't get tired. But a big question hangs over the classroom: Is the AI any better at separating the "smart ideas" from the "bad writing" than a human teacher is? If the AI is just as easily fooled by poor grammar as a tired teacher, then we might be building a grading system that unfairly punishes students who are still learning the language of science, mistaking their silence or stuttering for a lack of intelligence. This is the puzzle a team of researchers in Germany decided to solve.

The researchers set up a massive experiment to see if AI could tell the difference between a student who doesn't understand physics and a student who understands physics but writes poorly. They gathered 116 essays from ninth-grade students in Germany. The students were asked to explain a spacewalk scenario: Why can't astronauts hear each other shouting in the vacuum of space, but they can hear each other if they press their helmets together?

First, human experts graded these essays twice. Once, they gave a score for "Conceptual Understanding" (Did they get the physics right?). Separately, they gave a score for "Linguistic Quality" (Was the German grammar and vocabulary good?). This created a perfect control group where the "smartness" and the "writing skill" were measured independently.

Then, they fed these essays into nine different AI "brains." Some were traditional machine learning models (the old-school AI that learns by reading thousands of graded papers), and others were Large Language Models (the fancy new AI like the ones behind chatbots that can reason and follow instructions). The AI's job was simple: look at the essay and guess the "Conceptual Understanding" score.

The results were a bit of a wake-up call. The AI was generally good at its job, agreeing with the human experts about 67% to 78% of the time. However, when the researchers looked closer at the mistakes, a clear pattern emerged. The AI had a systematic blind spot. When a student wrote an explanation with lower linguistic quality (clunkier sentences, simpler words), the AI was significantly more likely to underestimate their understanding. In other words, the AI looked at a student with a great idea but messy writing and thought, "This student doesn't know physics," giving them a lower score than the human expert did.

Interestingly, the opposite didn't happen as often. If a student wrote beautifully but had weak physics ideas, the AI didn't consistently overestimate their score. The bias was specifically a penalty for poor writing. This happened across the board, whether the AI was a traditional machine learning model or a cutting-edge Large Language Model. It didn't matter which "brain" they used; they all seemed to struggle to look past the messy words to find the smart ideas.

The researchers found that for every step down in writing quality, the odds of the AI underestimating the student's physics knowledge jumped by about 1.3 to nearly 4 times. This suggests that the problem isn't just a glitch in one specific type of AI code. Instead, it seems to be a fundamental difficulty in the task itself. Because we can only see a student's thoughts through their language, it is incredibly hard—even for a super-smart computer—to tell if the language is bad because the ideas are bad, or if the language is bad just because the writer is still learning how to express themselves.

The study concludes that while AI is a powerful tool, it currently carries the same "language bias" that human teachers have struggled with for years. It tends to mistake a lack of vocabulary for a lack of understanding. This is a serious concern, especially for students who are learning physics in a language that isn't their first one. If schools start using these AI systems to make high-stakes decisions (like final exam grades) without fixing this bias, they risk systematically penalizing multilingual learners, misreading their struggle with words as a struggle with science. The paper suggests that until we can build AI that is truly robust to these linguistic differences, we need to be very careful about how we use it, perhaps using it as a helper rather than a final judge.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →