← Latest papers
💻 computer science

Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus

This paper benchmarks ten LLMs against Brazil's ENEM exam to reveal that while they can moderately rank question difficulty, they systematically underestimate item hardness, struggle with multimodal content, and lack the contextual plasticity needed for personalized assessment, suggesting they are best suited as calibrated screeners rather than authoritative generators.

Original authors: Thiago Brant, Julien Kühn, Jun Pang

Published 2026-02-09
📖 5 min read🧠 Deep dive

Original authors: Thiago Brant, Julien Kühn, Jun Pang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher trying to build a new math test. You have a super-smart robot assistant (an AI) that can write questions for you. But before you give the test to students, you need to know: Is this question too easy? Is it too hard?

Traditionally, teachers answer this by giving the test to thousands of students first, grading the results, and doing complex math to figure out the difficulty. This takes a long time and costs a lot of money.

This paper asks a simple question: Can we just ask the AI, "How hard is this question?" and trust its answer?

The researchers tested this idea using Brazil's massive national high school exam (ENEM), which has over 1,000 real questions. They asked ten different AI models to guess the difficulty of these questions and compared their guesses to the "official" difficulty scores calculated by human experts.

Here is what they found, explained with some everyday analogies:

1. The "Good at Ranking, Bad at Numbers" Problem

The AI models are like a student who is great at ordering books by size but terrible at measuring them with a ruler.

  • The Good News: If you ask the AI to sort 10 questions from "easiest" to "hardest," it does a decent job. It can tell you that Question A is harder than Question B.
  • The Bad News: If you ask the AI, "On a scale of 1 to 10, how hard is this?", it is consistently wrong. It almost always thinks the questions are easier than they actually are. It's like a person who thinks a marathon is just a "light jog" because they've never run one.

2. The "Blindfold" Effect (Visuals vs. Text)

Many exam questions have pictures, graphs, or charts. Since some AIs can't "see" images, the researchers had to describe the pictures in words (like a blind person describing a painting to a friend).

  • The Result: When the AI had to guess the difficulty based only on the text description of a picture, it got confused. It was much worse at judging these questions than the ones that were just text.
  • The Analogy: Imagine trying to guess how hard a puzzle is by reading a description of the pieces, rather than looking at the puzzle itself. You might miss a crucial clue, making the puzzle seem easier or harder than it really is.

3. The "One-Size-Fits-All" Trap (Student Backgrounds)

The researchers wanted to see if the AI could adjust its answer based on who is taking the test. They asked the AI: "How hard is this for a student from Finland?" vs. "How hard is this for a student from Brazil?"

  • The Result: The AI barely changed its mind. It was mostly "tone-deaf" to these changes.
  • The Analogy: Imagine a chef who cooks a soup. If you ask, "Is this soup too salty for a baby?" or "Is it too salty for a bodybuilder?", the chef just shrugs and says, "It's the same soup." The AI didn't really understand that different students have different backgrounds and knowledge levels. It couldn't adapt its "difficulty meter" to fit the person.

4. The "Prompt" Game (How you ask matters)

The researchers tried asking the AI in different ways. Sometimes they just said, "Rate this." Other times, they said, "Act like a math teacher, think step-by-step, and then rate this."

  • The Result: The way you ask the question changed the answer significantly. Some "thinking" prompts helped the AI get the ranking right, but they didn't fix the problem of it thinking everything was too easy.
  • The Analogy: It's like asking a friend for a movie recommendation. If you ask, "What's a good movie?", they might say a comedy. If you ask, "What's a good movie for a rainy Tuesday night?", they might say a thriller. The AI's answer depends heavily on how you phrase the question.

The Big Conclusion: The "Screening" Tool

The paper concludes that we shouldn't treat these AI models as Oracles (all-knowing gods who give the perfect answer). Instead, we should treat them as Screeners (a first filter).

  • The Recommendation: Use the AI to quickly scan a pile of new questions and say, "Hey, this one looks like it might be too easy, and this one looks like it might be too hard."
  • The Catch: You cannot trust the AI's exact number. You must fix its "bias" (teach it that it's underestimating difficulty) and you must have a human check the questions, especially the ones with pictures.

In short: AI is a helpful assistant that can help you sort the "easy" questions from the "hard" ones, but it is not ready to be the final judge of exactly how hard a question is, nor is it ready to understand the specific needs of different types of students. We need to calibrate it and keep a human in the loop.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →