← Latest papers
🔬 physics

Testing the Validity of Embedding-Based Similarity and Clustering for Handwritten Physics Solutions

This study demonstrates that while embedding-based similarity and clustering can support exploratory organization of handwritten physics solutions, they fail to reliably replicate human grading because the models prioritize surface features over conceptual understanding, necessitating external validation for assessment purposes.

Original authors: Maike Tauschhuber, Gerd Kortemeyer

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Maike Tauschhuber, Gerd Kortemeyer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher with a mountain of handwritten physics exams. You need to grade them, but there are hundreds of students, and reading every single one takes forever. You hear about a new technology called "AI embeddings" that can turn these handwritten answers into digital points on a map. The idea is simple: if two answers are close together on this map, they must be similar, right? So, maybe the AI can group similar answers together, saving you time.

This paper is like a reality check. The researchers asked: "Does this digital map actually reflect what a human teacher thinks is a 'good' answer?"

Here is what they found, explained through simple analogies:

1. The Experiment: The "Map" vs. The "Rubric"

The researchers took 992 handwritten thermodynamics exam answers from a tough engineering class. They used AI to:

  • Transcribe the handwriting into text in five different ways (some kept the messy handwriting style, some turned it into neat stories, some turned it into structured lists).
  • Map these texts onto a digital "embedding space" using nine different AI models.
  • Compare the AI's map to the actual scores given by human teachers.

Think of the human score as the "truth" (like a gold standard). The AI map is a guess at what that truth looks like.

2. The Result: The AI is a "Novice," Not an "Expert"

The researchers found that the AI map was not a perfect mirror of the human scores.

  • The "Surface Feature" Trap: The AI acted a bit like a physics novice. In physics education, there's a famous idea: Experts categorize problems by the deep physics principles (e.g., "This is about energy conservation"), while novices categorize them by surface features (e.g., "This problem has a lot of numbers" or "This one uses the letter 'Q'").

    • The AI embeddings behaved like the novices. They grouped answers together because they looked similar on the surface (same words, same symbols), even if the physics was wrong.
    • Conversely, they sometimes separated answers that were actually correct but written differently.
  • The "Fuzzy Neighborhood" Analogy: Imagine the AI creates neighborhoods on a map.

    • What the researchers hoped for: A neighborhood where everyone has an "A" grade, and another where everyone has an "F."
    • What they actually got: A neighborhood where you have a mix of "A"s, "B"s, and "C"s all living together. The AI said, "These answers are neighbors!" but the human teacher said, "No, this one is an A, and that one is a C."

3. The "Toy" Test: Why it Happened

To prove this, the researchers created a fake set of answers about a "Carnot engine" (a type of heat engine).

  • They wrote one correct answer.
  • Then, they wrote three incorrect answers that looked almost exactly the same, just with one tiny, fatal mistake (like using the wrong temperature unit).
  • The Result: The AI put the correct answer and the three wrong answers right next to each other on the map. It was fooled by the similar wording. It didn't "see" the deep physics error; it only saw the similar surface text.

4. What This Means for Grading

The paper concludes with a clear "Yes, but..."

  • The Good News: The AI isn't random. If two answers are close on the map, they do tend to have somewhat similar scores. The map is useful for exploration. You could use it to:

    • Find a batch of answers that look similar so a human can review them together.
    • Find "outliers" (answers that are very different from the rest) to check for unusual mistakes.
    • Help a teacher organize a pile of papers before grading.
  • The Bad News: The AI cannot grade the papers for you.

    • You cannot trust the AI to say, "These two answers are in the same cluster, so they get the same score."
    • The "distance" on the AI map does not equal the "distance" in grading points. A small shift on the map might mean a huge difference in points (like a sign error), or a huge shift might mean nothing.

The Bottom Line

Think of the AI embedding as a very helpful librarian, not a judge.

  • The librarian is great at saying, "Hey, these three books are on the same shelf and look similar."
  • But the librarian is not qualified to say, "Because these books are on the same shelf, they are all equally good."

The researchers say: Use the AI to organize the messy pile and help humans work faster, but never let the AI decide the final grade without a human looking at it first. The "geometry" of the AI map is enriched with some grading clues, but it is not a replacement for the human judgment required to understand deep physics concepts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →