Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment
This paper introduces a multimodal dataset comprising 485 scanned examination answers from 415 students across four institutions, paired with expert-designed Outcome-Based Education rubrics and comprehensive metadata to support research in automated evaluation, multimodal document understanding, and privacy-aware assessment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of education, the final grade on a test is often just a number, a single score that summarizes weeks of learning. But behind that number lies a complex process where a teacher reads a student's handwriting, interprets a sketch, checks a calculation, and decides how well a specific idea was understood. This process is known as outcome-based assessment, a method where the focus is not just on getting the right answer, but on demonstrating specific skills and knowledge steps along the way. For decades, computers have struggled to understand this human side of grading. They are excellent at reading typed text, but they often stumble when faced with the messy reality of a handwritten exam page filled with crossed-out words, scribbled diagrams, and equations written in the margins. To teach a computer to grade like a human, researchers need a massive library of real examples, paired with the detailed notes teachers use to decide the score. Without this, artificial intelligence remains blind to the nuances of how students actually show what they know.
A team of researchers at the Bangladesh Army University of Science and Technology has built exactly this kind of library. They have created a new collection of 485 scanned examination papers, gathered from 415 students across four different institutions. These are not simple text files; they are digital photographs of real exam scripts, complete with the original handwriting, printed diagrams, tables, and code. What makes this collection unique is that every single scanned page is linked to a detailed digital record that explains exactly how a human expert graded it. This record includes the original question, a model answer, and a specific set of rules called a rubric. A rubric breaks down a question into smaller parts, such as "understanding the concept" or "clarity of presentation," and assigns a score to each part based on how well the student performed. This allows the data to show not just that a student got 11 out of 15 points, but exactly which parts of their answer earned those points and which parts were missing.
The researchers gathered these materials from eight different faculty members who taught nine distinct subjects, ranging from computer science topics like machine learning and algorithms to more specialized fields like fisheries and e-commerce. The collection covers 12 different types of questions, which together contain 47 specific criteria for evaluation. To ensure the data is useful for training computers to handle real-world conditions, the team did not just scan the papers in a perfect, uniform way. They used a mix of mobile phone apps and traditional flatbed scanners. This means the images in the collection vary in brightness, contrast, and angle, just like real documents do in a classroom. Some pages are slightly tilted, some have shadows, and others are compressed differently. The students' work also varies widely; some papers are neat, while others contain crossed-out mistakes, inserted corrections, and arrows pointing to specific details. This variety is intentional, designed to help researchers build systems that can read and understand documents even when they are messy or imperfect.
The core of this work is the structure of the data itself. For every scanned PDF file, there is a corresponding digital file that acts as a map. This map links the image to the question asked, the correct answer, and a list of the specific criteria the teacher used to grade it. For each criterion, the map records how many points the student earned and describes what a perfect, good, average, or poor performance looks like for that specific part of the question. This level of detail is crucial because it moves beyond a simple total score. It allows researchers to train computers to identify exactly which parts of a handwritten answer are correct and which are not, rather than just guessing a final number. The team spent considerable time cleaning and organizing this data. They checked every score to make sure the math added up, standardized the names of the subjects, and replaced all student names and IDs with random codes to protect privacy. They also ensured that every file name matched its correct image, creating a secure and reliable dataset.
This collection is now available for other researchers to use, but with strict rules to protect the students involved. Because the images still contain the original handwriting and sometimes visible names or logos, the data is not open to the public. It is reserved for approved research projects where the goal is to improve how computers understand documents and assessments. The researchers emphasize that this data is a tool for study, not a system for grading real students. The scores in the collection were assigned by human teachers for their own classes, and the dataset is meant to help build better tools for the future, not to determine a student's actual grade today. By providing a large, diverse, and carefully labeled set of real exam papers, this work offers a solid foundation for the next generation of educational technology, helping to bridge the gap between human judgment and machine understanding.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.