Calibrating Gemini to Human Raters in Spoken Language Assessment for EFL Learners in Asia
This study demonstrates that while Gemini 2.5 Flash can achieve human-level agreement in assessing EFL learners' spoken English through 10-shot prompting when human raters show moderate consistency, its performance deteriorates when human agreement is excellent, highlighting the need for careful calibration and human supervision despite its ability to evaluate beyond mere word accuracy.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a teacher grading hundreds of essays. It's a mountain of work, and while you can easily read a student's written words, grading their spoken voice is a different beast entirely. Speech is "loud and transient"—it happens in a flash and then vanishes, making it hard to capture and judge fairly. For years, computers have been great at reading text, but they've struggled to "listen" to human conversation, often relying on written transcripts that miss the nuance of tone, pauses, and accent. Now, a new generation of "generative AI" has arrived. Think of these AIs as super-smart digital assistants that can not only read but also hear and understand context. The big question on everyone's mind is: Can we train these digital assistants to grade spoken English as fairly and accurately as a human teacher? This isn't just about saving teachers time; it's about ensuring that students get the right feedback to improve, regardless of where they come from or how they sound.
This paper takes a deep dive into that question, testing a specific AI model called Gemini 2.5 Flash. The researchers wanted to see if they could "calibrate" this AI—essentially, teach it how to grade—by showing it examples of how human teachers score spoken dialogues. They didn't just feed the AI written text; they gave it the actual audio files of students speaking, which is a big step forward because it lets the AI hear the real voice, not just a transcription.
The researchers set up a clever experiment using 129 recorded conversations between Asian English learners and their teachers. They asked the AI to grade these speeches using three different methods:
- Zero-shot: The AI just got the rules and graded immediately, with no examples.
- 5-shot: The AI saw five examples of how a human teacher graded before it started.
- 10-shot: The AI saw ten examples.
They compared the AI's scores against two pairs of human teachers. One pair of teachers agreed with each other moderately well (their scores were similar, but not identical), while another pair agreed almost perfectly (they were like twins in their grading).
Here is the twist: The AI worked best when it was learning from the teachers who had moderate agreement. When the AI was shown examples from the "perfectly agreeing" teacher pair, it actually got worse, failing to match their scores at all. It's as if the AI needed to see a little bit of human disagreement to understand the full range of what a "good" or "bad" score looks like. If the examples were too perfect, the AI got confused about where the boundaries were.
The study also checked if the AI could actually "hear" the students correctly. They measured this using something called Word Error Rate (WER), which counts how many words the AI got wrong in its transcription. The AI made mistakes, but mostly it was just mixing up who was speaking (the student or the teacher) rather than missing the words entirely. The researchers found that the AI's ability to hear the words didn't actually change how harshly or kindly it graded the students. This suggests the AI was listening to more than just the dictionary definitions of the words; it was picking up on the flow and rhythm of the speech, too.
However, there was a small, surprising snag. The AI seemed to give slightly lower scores to students from Pakistan, even though the AI transcribed their speech very accurately. The researchers noted this as a potential bias that needs more investigation, especially since there were only four students from that country in the study.
In the end, the paper suggests that Gemini 2.5 Flash is a promising tool for helping teachers grade spoken English, but it's not a magic wand. It works best when it's calibrated with human examples that show a realistic range of scores, and it needs a human supervisor to keep an eye on it. The AI can hear the speech, understand the context, and give feedback that looks a lot like a human teacher's, but it still needs careful tuning to make sure it's fair to everyone.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.