← Latest papers
💬 NLP

A Survey of Automated Presentation Coaching: Systems, Methods, and Open Challenges

This paper presents the first systematic survey of automated presentation coaching systems, introducing a five-dimensional task taxonomy to map existing tools across pronunciation, prosody, and content domains while identifying key technical methods and critical open challenges such as data scarcity and accent fairness.

Original authors: Wen Liang, Li Siyan, Zackary Rackauckas, Julia Hirschberg

Published 2026-06-29
📖 5 min read🧠 Deep dive

Original authors: Wen Liang, Li Siyan, Zackary Rackauckas, Julia Hirschberg

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are preparing for a big speech, like a TED Talk or a company presentation. You know your content, but you're worried about how you sound: Are you speaking too fast? Do you stress the wrong syllables? Does your voice sound flat?

This paper is a comprehensive map of the "digital coaches" currently available to help you practice. The authors, researchers from Columbia University and others, looked at 15 different computer systems designed to give you feedback on your speaking. They didn't just list them; they built a new way to measure what these systems can and cannot do.

Here is a simple breakdown of their findings, using some everyday analogies.

1. The "Five-Point Report Card"

The authors realized that previous research treated "speaking well" as a single thing. Instead, they created a five-dimensional checklist (a taxonomy) to grade any coaching system. Think of this like a report card for a speech:

  1. Pronunciation (The Letters): Are you saying the individual sounds and words correctly? (e.g., saying "algorithm" instead of "al-gor-ith-m").
  2. Lexical Stress (The Emphasis): Are you hitting the right beat in multi-syllable words? (e.g., stressing the first syllable in "AL-go-rithm" rather than the second).
  3. Prosody (The Music): Is your voice rising and falling naturally to show excitement, questions, or pauses? This is the "melody" of your speech.
  4. Pacing (The Rhythm): Are you speaking at a good speed? Are you pausing in the right places, like when you switch slides?
  5. Content Faithfulness (The Script): Did you actually say what you were supposed to say? Did you skip important technical terms or add random words?

2. The Current State of the Field: "The Patchwork Quilt"

The authors surveyed existing systems and found that they are like a patchwork quilt where some squares are very detailed, but others are missing entirely.

  • What they do well: Most systems are great at Pronunciation (checking if you said the right letters) and Pacing (telling you if you are too fast). It's like having a coach who is excellent at checking your spelling and your watch.
  • What they ignore:
    • Lexical Stress is almost completely ignored. The paper notes that while this is crucial for being understood, very few systems check if you are stressing the right part of a word.
    • Content Faithfulness is also rare. Most systems don't check if you actually mentioned the specific keywords from your slides.
  • The Missing Piece: No single system currently checks all five of these areas at once. The most advanced system mentioned, PresentCoach, covers four of them, but it still misses the "Stress" dimension.

3. How These Coaches Work: The "Shadowing" Method

The paper explains that these systems generally use two main tricks, similar to how a music student learns a song:

  • The "Perfect Model" (TTS): The computer uses AI to generate a perfect version of what you should sound like. It can even mimic your own voice but with better rhythm and stress. This is like a music teacher playing the song perfectly so you can copy it.
  • The "Side-by-Side" Comparison: The system records you and lines it up with the "Perfect Model." It then highlights exactly where you drifted off.
    • Analogy: Imagine a dance instructor recording your routine and playing it back next to the champion's routine. The computer highlights, "You missed a step here," or "You held that pose too long."

4. The Big Problems (The "Open Challenges")

Even though the technology is impressive, the authors point out three major hurdles that stop these systems from being perfect:

  • The "Empty Library" Problem: To teach a computer what a good presentation sounds like, you need a massive library of recordings from non-native speakers giving actual presentations with slides. The paper says this library doesn't exist yet. Most available data is just people reading short sentences in a quiet room, not giving a 20-minute talk.
  • The "Accent Bias" Problem: Current systems often treat a specific accent as "wrong." The authors argue that a good coach should help you be clear without forcing you to lose your unique accent identity. It's like a coach helping a runner improve their form, not trying to make them run like someone from a different country.
  • The "Real-Time" Gap: Most systems wait until you finish your whole speech to give you feedback. This is like getting a report card at the end of the semester. The authors say we need systems that can whisper advice while you are practicing, but doing this instantly on a phone or laptop is still very hard to build.

5. The Bottom Line

The paper concludes that while we have great tools for checking individual parts of a speech (like spelling or speed), we don't have a "super-coach" yet that handles the whole performance—checking your stress, your rhythm, your content, and your accent fairness all at once.

The authors are calling on the research community to build better datasets (more real-world speech recordings) and to create systems that can give instant, fair feedback to speakers from all over the world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →