← Latest papers
💻 computer science

SCALEFeedback: A Large-Scale Dataset of Synthetic Computer Science Assignments for LLM-generated Educational Feedback Research

This paper introduces SCALEFeedback, a large-scale open-source dataset of 10,000 synthetic computer science assignments and student submissions generated via the Sophisticated Assignment Mimicry (SAM) framework to advance research on scalable, effective, and responsible LLM-based educational feedback.

Original authors: Keyang Qian, Kaixun Yang, Wei Dai, Flora Jin, Yixin Cheng, Rui Guan, Sadia Nawaz, Zachari Swiecki, Guanliang Chen, Lixiang Yan, Dragan Gašević

Published 2026-07-03
📖 5 min read🧠 Deep dive

Original authors: Keyang Qian, Kaixun Yang, Wei Dai, Flora Jin, Yixin Cheng, Rui Guan, Sadia Nawaz, Zachari Swiecki, Guanliang Chen, Lixiang Yan, Dragan Gašević

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a university classroom where a teacher has 1,000 students. Each student turns in a homework assignment, and the teacher needs to write a personalized, helpful comment on every single one. In the real world, this is a nightmare. There aren't enough hours in the day, and the teacher gets tired.

Enter Artificial Intelligence (AI). We want to teach the AI to be the "super-teacher" that writes these comments instantly. But there's a huge problem: AI needs to practice. To learn how to give good feedback, the AI needs to see real student homework and real teacher comments.

However, schools can't just hand over real student papers to the public. That would be a massive privacy violation (imagine your name and bad grades being posted online) and a copyright nightmare.

The Solution: A "Digital Twin" Classroom
This paper introduces a clever solution called SCALEFeedback. Think of it as building a perfectly realistic, fake classroom that looks and feels exactly like the real one, but where every single student is a robot, and every paper is made of pure imagination.

Here is how they built it, using a simple analogy:

1. The "Sophisticated Assignment Mimicry" (SAM) Framework

Imagine you are an actor trying to play a specific character. You don't just guess what they say; you study them closely.

  • The Real World: The researchers took real university assignments (the "script") and real student submissions (the "performance").
  • The Mimicry: They used a super-smart AI (an LLM) to act as a "digital twin."
    • Step 1 (The Study): The AI looked at a real assignment and a real student's answer. It analyzed everything: the tone, the length, the mistakes made, and the grade given.
    • Step 2 (The Performance): The AI then wrote a brand new assignment and a brand new student answer that sounded exactly like the real ones but used completely made-up words.
    • Step 3 (The Double-Check): The AI reviewed its own fake work. "Does this look like the real thing? Did I accidentally copy a real student's name?" If it failed the check, it started over.

This process was repeated 10,000 times. The result is a dataset of 10,000 fake student papers from 59 different computer science courses.

2. Why is this dataset special?

Usually, when people make fake data, it's like a cartoon drawing of a house—it looks like a house, but it's not very detailed.

  • The "Naïve" Way: If you just ask an AI to "make up a homework assignment," it might write something generic. It's like a bad photocopy.
  • The SAM Way: This paper claims their method is like a high-definition 3D scan. They checked the fake data against the real data and found:
    • The "Vibe" is the same: The fake papers use the same types of words and sentence structures as the real ones.
    • The Grades match: The fake students got grades that looked just like the real students' grades.
    • The Length is right: The fake papers are the same length as the real ones.

3. Did it work? (The "Taste Test")

To prove this fake classroom was useful, the researchers did a "taste test."

  • They took the Real Papers and asked 10 different AI models to give feedback.
  • They took the Fake Papers and asked the same 10 AI models to give feedback.
  • The Result: The feedback the AI gave to the fake papers was almost identical to the feedback it gave to the real papers.

The Analogy: Imagine a chef tasting a soup made with real chicken and a soup made with a perfect plant-based substitute. If the chef can't tell the difference and says, "Both taste great," then the substitute is a success. The researchers found that the AI "taste testers" couldn't tell the difference between the real and fake student work.

4. The Safety Net (Privacy)

The most important rule was: No real names allowed.
The researchers built a "Privacy Gate." Before any fake paper was saved, a super-smart AI guard checked it.

  • The Guard's Job: "Did you accidentally copy a real student's name or ID number?"
  • The Result: If the guard found a real name, the paper was thrown in the trash and rewritten. The final dataset is 100% clean of private information.

Summary of the Paper's Claims

  • What they made: A free, open library of 10,000 fake computer science assignments and student answers.
  • How they made it: Using a "mimicry" process where AI copies the style and structure of real data without copying the actual content.
  • What they proved:
    1. The fake data looks statistically identical to real data (same length, same grades, same word choices).
    2. AI models give the same quality of feedback to the fake data as they do to real data.
    3. The data is safe; no real student privacy was leaked.
  • What they didn't claim: They did not claim this dataset is perfect for everything. They admit that the fake data is a bit "too average." It doesn't capture the extreme outliers (like a student who wrote a 50-page essay or a student who failed spectacularly). It's great for general training, but maybe not for training AI to spot extreme edge cases.

In a nutshell: This paper gives researchers a safe, legal, and realistic "playground" to teach AI how to be a better teacher, without ever risking a single student's privacy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →