← Latest papers
💬 NLP

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track

NAVER LABS Europe's IWSLT 2026 submission achieves a tied first-place ranking in the instruction-following short track by enhancing its multi-stage ASR, ST, and SQA pipeline with the SpeechMapper projector and a synthetic scientific dataset (fakACL), enabling superior performance with a more compact architecture and a weaker LLM backbone.

Original authors: Marcely Zanon Boito, Hemant Yadav, Jean-Luc Meunier, Ioan Calapodescu

Published 2026-07-03
📖 4 min read☕ Coffee break read

Original authors: Marcely Zanon Boito, Hemant Yadav, Jean-Luc Meunier, Ioan Calapodescu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a very smart, but slightly forgetful, robot assistant how to listen to a human speak, understand what they said, and then answer questions about it—all in different languages. This is exactly what the team at NAVER LABS Europe did for a major competition called IWSLT 2026.

Here is a breakdown of their work, explained simply with some analogies.

The Goal: The "Universal Translator" Challenge

The team entered a contest where they had to build a system that could:

  1. Listen to English speech and write it down (like a court stenographer).
  2. Translate that speech into Chinese, Italian, or German.
  3. Answer questions about what was heard, in those same languages.

Think of this as training a robot to sit in a lecture hall, take notes, translate the lecture for a friend, and then quiz the friend on the content, all without ever seeing the text on a screen—only hearing the voice.

The Problem: A Smaller Brain, A Bigger Job

Last year, this team won the competition using a massive "brain" (a Large Language Model or LLM). This year, they were forced to use a much smaller, weaker brain (a 4-billion parameter model).

The Analogy: Imagine last year's robot had a PhD in linguistics. This year, they had to use a very bright high school student. The challenge was to make this "student" perform as well as the "PhD" without giving them extra time or resources.

The Solution: Two Specialized Tools

To make this smaller brain work, they didn't just throw data at it. They built two specialized tools to help the brain learn, like a tutor and a translator working together.

1. The "Ear-to-Brain" Connector (SpeechMapper)

The robot's brain is designed to read text, not hear sound. Usually, you need a heavy, complex machine to turn sound waves into text for the brain to understand.

  • What they did: They replaced the old heavy machine with a new, lighter tool called SpeechMapper.
  • The Analogy: Imagine the old method was like trying to translate a song by writing down every single note on a sheet of music before singing it. It was slow and clunky. The new SpeechMapper is like a "musical ear" that instantly recognizes the melody and hums the right tune directly into the singer's ear. It learns to turn sound into "thoughts" using only listening practice, making it much faster and less demanding on the computer's memory.

2. The "Fake Lecture" Generator (fakACL)

The team noticed that the robot was good at reading books but struggled with the specific style of the test: scientific presentations (like those at a computer science conference).

  • What they did: They created a fake dataset called fakACL. They used the robot's own brain to write fake scientific speeches, then used a voice synthesizer to turn those speeches into audio.
  • The Analogy: It's like a student who is great at studying history textbooks but is about to take a test on modern physics. Instead of just hoping they pass, the teacher (the team) generates a bunch of fake physics lectures for the student to practice on. This bridges the gap between what the student knows and what the test asks.

The Training Process: Parallel Learning

The team didn't train everything at once. They used a three-step pipeline:

  1. Step A: Teach the "Ear" (SpeechMapper) to understand sound using real speech data.
  2. Step B: Teach the "Brain" (the LLM) to understand text and answer questions using text data.
  3. Step C: Bring them together. They combined the trained "Ear" and the trained "Brain" and gave them a final, short practice session where they had to listen and speak at the same time.

The Results: Punching Above Their Weight

Despite using a much smaller "brain" than last year, their new system performed incredibly well.

  • The Outcome: They tied for first place in the overall rankings.
  • The Surprise: Their system was actually better than last year's winning system in some areas, even though it was smaller and used less computing power.

Why It Matters (According to the Paper)

The paper claims that by using a smarter way to connect sound to text (SpeechMapper) and by practicing on realistic, fake data (fakACL), they proved you don't need a giant, expensive computer brain to build a great voice assistant. You just need the right training methods.

In short: They took a smaller, cheaper robot, gave it a better "ear," practiced it on fake lectures, and it ended up winning the race against the giants.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →