← Latest papers
🤖 AI

SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

This paper introduces SEFORA, a large-scale corpus of student essays with real instructor feedback, and UniMatch, a novel evaluation framework that reveals current LLMs struggle to generate feedback aligning with human instructor priorities, achieving a maximum F1 score of only 0.4 across extensive experiments.

Original authors: Shayan Peyghambari Oskoui, Norah Almousa, Zhaoyi Joey Hou, Carolina Gustafson, Gayle Rogers, Raquel Coelho, Diane Litman, Xiang Lorraine Li

Published 2026-07-02
📖 4 min read☕ Coffee break read

Original authors: Shayan Peyghambari Oskoui, Norah Almousa, Zhaoyi Joey Hou, Carolina Gustafson, Gayle Rogers, Raquel Coelho, Diane Litman, Xiang Lorraine Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher grading a stack of student essays. You have to read every sentence, find the good parts, point out the confusing parts, and write specific notes in the margins to help the student improve. It's hard work, and doing it for hundreds of students at once is nearly impossible for a human.

Recently, we've tried to use "AI teachers" (Large Language Models) to do this job. But there's a problem: we don't really know if the AI is giving good advice, and we don't have a good ruler to measure it.

This paper introduces two things to fix that: a massive library of real teacher notes called SEFORA, and a new measuring tool called UNIMATCH.

1. The Library: SEFORA (The "Gold Standard" Collection)

Think of SEFORA as a giant, organized scrapbook.

  • What's inside: It contains 564 drafts of student essays (some written once, some rewritten multiple times) from real college classes.
  • The special part: Next to every student essay, it has the actual notes written by real human professors. These aren't just red marks saying "wrong." They are sticky notes and highlights that say things like, "This character feels real," or "What does 'that' refer to here?"
  • Why it matters: Before this, most AI datasets only had scores (like "85/100") or simple error tags. SEFORA is special because it captures the nuance of how real teachers talk to students, including the specific parts of the text they are talking about.

2. The Ruler: UNIMATCH (The "Feedback Detective")

Now, imagine you ask an AI to write feedback on a student's essay. How do you know if it's any good?

  • The old way: You might compare the AI's words to the teacher's words to see if they look similar. But this is like judging a chef by whether they used the same words as the recipe, rather than if the food tastes right. If the teacher says "Make it shorter" and the AI says "Trim the fat," a simple word-count check might say they are different, even though they mean the same thing.
  • The new way (UNIMATCH): This tool acts like a detective. It breaks both the teacher's notes and the AI's notes down into tiny, individual "feedback units" (one specific thought per unit).
    • Step 1: It splits the notes into pieces.
    • Step 2: It asks, "Does this piece from the AI match the meaning of a piece from the teacher?" (e.g., Does "Trim the fat" match "Make it shorter"?).
    • Step 3: It uses a smart matching system (like pairing socks) to see how many of the teacher's important points the AI managed to find and how many extra, useless points the AI invented.

3. The Big Discovery: The "Chatterbox" Problem

The researchers tested 74 different ways of asking AI models to write feedback. The results were surprising and a bit disappointing:

  • The Score: Even the smartest AI models only managed to get a score of about 0.4 out of 1.0. This means they are missing most of the specific, high-priority feedback a human teacher would give.
  • The Main Culprit: The biggest reason for the low scores was verbosity (talking too much).
    • The Analogy: Imagine a student asks for help with one math problem. A good tutor gives one clear hint. The AI, however, acts like a nervous chatterbox. It gives the one hint, but then adds five more suggestions that the teacher never would have made, or repeats the same point three different ways.
    • The Result: Because the AI generates so much "extra" noise, its precision drops. It's like a student who writes a 10-page essay to answer a one-sentence question; they might get the right answer, but they've buried it under so much fluff that it's not helpful.

Summary

The paper tells us that while AI can generate text, it currently struggles to act like a thoughtful human teacher. It tends to over-explain and miss the specific, high-value points that real instructors prioritize.

To fix this, we need:

  1. Better Data: Real examples of how teachers actually talk (which SEFORA provides).
  2. Better Measurement: A way to judge if the AI is hitting the right points, not just using the right words (which UNIMATCH provides).

Until AI learns to be concise and hit the right targets, it's not quite ready to replace the human touch in the classroom.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →