← Latest papers
💬 NLP

PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech

This paper introduces INSV-A, an automated screening framework for low-resource non-Latin-script text-to-speech evaluation, and instantiates it as PashtoTTS-Bench to systematically assess multiple TTS systems on Pashto using metrics like ASR word error rate, script fidelity, and language identification.

Original authors: Hanif Rahman

Published 2026-05-27
📖 5 min read🧠 Deep dive

Original authors: Hanif Rahman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a judge at a talent show for robots that can talk. Your job is to test if they can speak Pashto, a language spoken by millions in Afghanistan and Pakistan that uses a unique script (Perso-Arabic) and has some very tricky sounds that don't exist in English or even Urdu.

For a long time, judges only had one way to grade these robots: they would listen to the robot, write down what they heard, and count how many words were wrong. This is called the "Word Error Rate" (WER).

The Problem:
This old method is like judging a chef who is trying to cook a spicy curry but accidentally serves you a bowl of plain rice. If the judge just counts the words, they might say, "Well, the rice is perfectly spelled, so the chef gets a good score!" But the judge missed the fact that the chef served the wrong dish entirely.

In the world of Pashto TTS (Text-to-Speech), robots often make three sneaky mistakes that a simple word-count misses:

  1. The Wrong Language: They read Pashto text but speak Urdu or Dari (neighboring languages).
  2. The Wrong Script: They speak the words but use the wrong letters in the transcript (like writing English letters instead of Pashto ones).
  3. The Robot Voice: They say the right words perfectly, but the voice sounds flat, mechanical, and unnatural to a human ear.

The New Solution: INSV-A

The author of this paper, Hanif Rahman, built a new "scorecard" called INSV-A. Instead of just one number, it's a four-part checklist that acts like a security scanner at an airport.

Here is how the four parts work, using simple analogies:

  1. Intelligibility (The "Can you hear me?" check):

    • The Test: Does the robot actually make sound? Can a computer transcribe the words?
    • The Analogy: This is checking if the radio is turned on and if the signal is clear enough to understand the words.
  2. Script Fidelity (The "Pen and Paper" check):

    • The Test: When the computer writes down what the robot said, does it use the correct Pashto letters?
    • The Analogy: If you ask someone to write a poem in Pashto, and they hand you a piece of paper with English letters, they failed the test, even if they said the poem correctly. This check ensures the "writing" matches the "language."
  3. Verification (The "Passport" check):

    • The Test: Is the robot actually speaking Pashto, or did it switch to Urdu?
    • The Analogy: This is like checking a passport at the border. The robot might look like it's speaking Pashto, but this check asks, "Are you really a Pashto speaker, or are you an imposter speaking a neighboring language?"
  4. Naturalness (The "Human Ear" check):

    • The Test: Does it sound like a real person, or a robot?
    • The Analogy: This is the "vibe" check. Even if the words are perfect, does it sound like a friendly neighbor or a cold machine?
    • Note: The paper says this part is planned but not fully done yet in this specific release. They set up the test, but they haven't finished grading the "vibe" scores.

The Results: The "PashtoTTS-Bench"

The author tested this new scorecard on several robots in April and May 2026. Here is what they found:

  • The "Automatic" Robot (OmniVoice auto): This robot was the surprise winner. It got the lowest number of word errors.
    • The Catch: It scored better than real human speakers. Why? Because the computer used to grade it is bad at understanding messy, real human speech (with accents and background noise), but it loves the clean, perfect sound of a robot. So, a low error score here just means the robot sounds "clean," not necessarily that it sounds "better than a human."
  • The "Commercial" Robots (Edge GulNawaz & Latifa): These are the official Microsoft voices. They did a solid job, speaking clear Pashto with the right letters. They are the "safe" baseline.
  • The "Cloned" Robot (OmniVoice clone): This robot tried to copy a voice but stumbled a bit more, failing to speak some sentences at all.
  • The "Urdu" Robot (The Control): The author tested a robot that should speak Urdu. It correctly failed the Pashto test, proving the new scorecard can tell the difference between Pashto and its neighbor, Urdu.

The Big Discovery: The "Whisper" Trap

One of the most important findings is about a popular tool called Whisper (a famous AI for speech).

  • The paper found that Whisper is blind to Pashto. Even though Pashto is in its dictionary, it almost never recognizes it. It thinks Pashto is something else.
  • The Lesson: You cannot rely on just one tool to judge these languages. You need a team of different tools (like MMS and SpeechBrain) to double-check each other, or you might miss the failure entirely.

Summary

This paper doesn't just give a score; it gives a new rulebook. It tells us that for languages like Pashto, you can't just count words. You have to check:

  1. Did it speak the right language?
  2. Did it use the right letters?
  3. Did it actually make sound?
  4. (Eventually) Did it sound human?

The author has released all the tools, the test questions, and the scoring scripts so that anyone can run this test again in the future to see if new robots get better. It's a "checkpoint" to make sure we aren't just building robots that look like they speak Pashto, but actually do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →