← Latest papers
⚡ electrical engineering

The NTNU System at the S&I Challenge 2025 SLA Open Track

The NTNU team's system for the S&I Challenge 2025 SLA Open Track integrates wav2vec 2.0 with the Phi-4 multimodal large language model via a score fusion strategy to overcome modality-specific limitations, achieving a second-place ranking with a root mean square error of 0.375.

Original authors: Hong-Yun Lin, Tien-Hong Lo, Yu-Hsuan Fang, Jhen-Ke Lin, Chung-Chun Wang, Hao-Chien Lu, Berlin Chen

Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Hong-Yun Lin, Tien-Hong Lo, Yu-Hsuan Fang, Jhen-Ke Lin, Chung-Chun Wang, Hao-Chien Lu, Berlin Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher trying to grade a student's spoken English. You have to listen to them and decide: How good are they? Do they sound fluent? Is their pronunciation clear? Do they use the right words?

For a long time, computers tried to do this by doing two separate things, but they were like two specialists who couldn't talk to each other:

  1. The "Transcriber" (BERT): This computer listens, writes down exactly what the student said, and then grades the words. It's great at checking if the student used the right vocabulary. But, if the computer mishears a word (which happens often with accents), the whole grade gets messed up. Also, it can't hear how the student said it—like if they sounded nervous or spoke in a weird rhythm.
  2. The "Sound Engineer" (wav2vec 2.0): This computer ignores the words entirely. It just listens to the sound waves. It's amazing at hearing if the student is speaking smoothly, if their accent is clear, and if they are pausing too much. But, it doesn't really understand what the student is saying. It might think a student is great just because they spoke fluently, even if they were saying nonsense.

The NTNU Team's Solution: The "Super-Teacher"

The team from National Taiwan Normal University (NTNU) realized that to get a perfect grade, you need both the Transcriber and the Sound Engineer working together. They built a system for the "Speak & Improve Challenge 2025" (a big competition for AI that grades speaking) that acts like a Super-Teacher.

Here is how their system works, using a simple analogy:

1. The Two Graders
They built two distinct "AI graders" that listen to the same student recording at the same time:

  • Grader A (The Sound Expert): Uses a model called wav2vec 2.0. It focuses entirely on the music of the speech—the rhythm, the pronunciation, and the flow.
  • Grader B (The Meaning Expert): Uses a massive, smart AI called Phi-4. This is a "Multimodal Large Language Model." It's like a super-smart tutor that can read the text and listen to the voice at the same time. It understands the meaning, the grammar, and the context of the story the student is telling.

2. The "Score Fusion" (The Final Decision)
Instead of letting one AI make the final call, they combined the two. Imagine a panel of judges where one judge votes on "Fluency" and the other votes on "Content."

  • They didn't just average the scores blindly. They used a smart strategy called Score Fusion.
  • They looked at different "levels" of English (from beginner to advanced). For each level, they figured out exactly how much weight to give the Sound Expert versus the Meaning Expert to get the most accurate result.
  • It's like saying, "For a beginner, let's trust the Sound Expert a little more. For an advanced student, let's trust the Meaning Expert more."

The Result: Second Place!

When they tested this "Super-Teacher" system against the official test set:

  • The Old Way (Baseline): Made a mistake of about 0.44 points on average.
  • The Sound Expert Alone: Made a mistake of 0.39.
  • The Meaning Expert Alone: Made a mistake of 0.39.
  • The NTNU Super-Teacher (Combined): Made a mistake of only 0.375.

This small improvement was enough to secure 2nd place in the entire competition, beating out many other top researchers. The only team that did slightly better (1st place) had a score of 0.364.

Why This Matters (According to the Paper)

The paper claims that this proves you can't just rely on one type of AI.

  • If you only look at the words, you miss the feel of the speech.
  • If you only listen to the sounds, you miss the meaning.
  • By combining a specialized "Sound Expert" with a "Super-Smart Tutor" and letting them vote together, you get a much fairer and more accurate grade.

The team also found that having a specific "Tutor" for each specific type of test question (like an interview vs. a presentation) worked better than having one "Tutor" try to answer everything at once.

In short: To grade a human speaking, you need a computer that can hear the music and understand the lyrics, working together as a team.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →