A Semi-Automated System for Generating Dialogue-Based TTS Lessons Using Large Language Models: An Exploratory Study of Educational Potential
This exploratory study demonstrates that a semi-automated system using Large Language Models to generate expert-novice dialogue-based Text-to-Speech lessons significantly enhances student comprehension and cognitive engagement compared to single-speaker TTS, offering a viable, educator-augmented alternative to traditional instructor voice despite a trade-off in audio naturalness.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: Semi-Automated Generation of Dialogue-Based TTS Lessons Using LLMs
Problem Statement
The production of high-quality educational content, particularly for flipped classroom models, demands significant time and specialized skills from educators, creating a high barrier to entry. While fully automated content generation using Generative AI and Text-to-Speech (TTS) has emerged, these solutions often lack pedagogical nuance, factual accuracy, and optimization for "listenability." Furthermore, LLM-generated content carries risks of hallucinations, necessitating human oversight. Existing automated systems typically focus on technical evaluation or simple read-alouds, failing to validate educational effectiveness in actual classroom settings or to leverage the unique capabilities of TTS for dialogue-based instruction.
Methodology
The study proposes and validates a semi-automated, human-in-the-loop system designed to augment, rather than replace, educators. The system operates through a three-stage workflow:
- Structured Slide Generation: Unstructured text (e.g., textbooks) is processed by an LLM (Claude 3.5 Sonnet) to generate Marp-format slides. Educators review these for fact-checking, pedagogical flow, and visual refinement.
- TTS-Optimized Narration Generation: The LLM generates narration scripts optimized for TTS "listenability." The study introduces a novel Expert-×-Novice dialogue format, inspired by cognitive apprenticeship theory. In this mode, the LLM creates a script where an "Expert" provides knowledge and a "Novice" asks questions, rephrases concepts, and confirms understanding. This structure is designed to scaffold learning and induce self-explanation effects. Educators review these scripts for accuracy and clarity.
- Speech Synthesis and Integration: Using the Gemini TTS API, the system synthesizes audio for the scripts. Two modes are supported: single-speaker (teacher-like) and dialogue (Expert and Novice voices). The audio is synchronized with slide images (rendered via Selenium) and integrated into final video files using MoviePy.
Empirical Validation
The system's educational potential was evaluated through a quasi-experiment involving 245 first-year high school students in Japan. Participants sequentially experienced three lesson formats:
- Instructor-voice video (Control).
- Single-speaker TTS lesson.
- Dialogue TTS lesson (Expert-×-Novice).
Note: Due to the fixed order and varying content across sessions, the study acknowledges constraints in separating format effects from content and order effects.
Data collection included retrospective within-subject comparisons (comparing all three formats) and repeated cross-sectional between-group comparisons (Single TTS vs. Dialogue TTS). Evaluation metrics were grounded in the ARCS model (motivation), Cognitive Load Theory (extraneous vs. germane load), and learning engagement theory.
Key Results
RQ1: Does TTS degrade the learning experience?
Within-subject analysis of core metrics (comprehension, concentration, overall evaluation) found no significant differences between instructor-voice, single TTS, and dialogue TTS. TOST equivalence testing (bounds ) confirmed that all metrics fell within the equivalence range. This suggests TTS audio does not undermine the fundamental learning experience compared to human voice.RQ2: Is the dialogue format educationally superior?
Comparing Single TTS vs. Dialogue TTS revealed a trade-off:- Advantages of Dialogue TTS: Significantly higher scores in comprehension () and cognitive engagement (). Learners felt more confident in their ability to explain the content to others. Supplementary analysis using a proportional odds model (controlling for prior knowledge) confirmed these advantages were not solely due to prior knowledge imbalances.
- Advantages of Single TTS: Significantly higher scores in audio naturalness (). The dialogue format introduced higher extraneous cognitive load, likely due to variations in audio quality between speakers and pages.
- Preference: Dialogue TTS was selected as the "most enjoyable to learn from" by 66.9% of participants, significantly outperforming other formats.
Key Contributions
- System Architecture: A practical, three-stage human-in-the-loop workflow that balances LLM automation with educator quality control, specifically designed to generate TTS-optimized scripts.
- Novel Methodology: The establishment of an automated Expert-×-Novice dialogue generation method inspired by cognitive apprenticeship theory, tailored for TTS listenability.
- Empirical Evidence: The first study to integrate LLM automation, TTS audio acceptability, and dialogue format design in a quasi-experimental setting with actual learners. It provides evidence that dialogue-based TTS can enhance comprehension and engagement despite potential audio quality trade-offs.
Significance and Claims
The paper claims to provide a theoretical and empirical basis for the educational acceptability of TTS audio and specific design principles for TTS lesson formats.
- It demonstrates that TTS audio does not inherently degrade learning compared to human voice, lowering the barrier for content production.
- It suggests that leveraging the flexibility of TTS to create dialogue-based lessons can improve self-assessed comprehension and cognitive engagement, provided the audio quality trade-offs are managed.
- The authors explicitly state that these findings are based on a quasi-experiment with fixed-order and varying-content constraints. Therefore, they do not claim definitive causal effects of the lesson format alone but rather offer design implications and patterns observed. They emphasize that future work requires randomized crossover designs and objective learning tests to confirm causal relationships.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.