← Latest papers
💻 computer science

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations

This paper presents Human-1, the first open and reproducible full-duplex spoken dialogue system for Hindi, which adapts the Moshi architecture using a custom tokenizer and 26,000 hours of real-world conversational data to enable natural speech behaviors like interruptions and overlaps.

Original authors: Bhaskar Singh, Shobhit Banga, Pranav Sharma

Published 2026-04-28
📖 4 min read☕ Coffee break read

Original authors: Bhaskar Singh, Shobhit Banga, Pranav Sharma

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The "Human-1" Project: Teaching AI to Master the Art of the "Chat"

Imagine you are talking to a very polite, but slightly robotic, assistant. You say, "Hey, can you tell me about..." and then you pause to think. The assistant just sits there in dead silence, waiting for you to finish your entire sentence perfectly. If you accidentally interrupt it, it gets confused and restarts. It’s a "walkie-talkie" style of talking: one person speaks, then the other speaks.

Now, imagine talking to your best friend. You both talk at the same time sometimes. You make little "mhm" or "yeah" sounds while they are talking to show you're listening. If they stop suddenly, you jump in. This is "Full-Duplex" conversation—it’s fluid, messy, and natural.

The Problem: Most AI models are "walkie-talkies." They are great at English, but they don't understand the "rhythm" of how people actually talk in other languages, like Hindi.

The Solution: A team at Josh Talks created Human-1, the first system designed to let an AI have a real, flowing, "human-style" conversation in Hindi.


How They Did It (The Three Big Ingredients)

1. The "Giant Library" of Real Talk (The Data)

Most AI is trained on books or news reports—things people read. But you don't talk like a textbook. To fix this, the researchers collected 26,000 hours of real, spontaneous Hindi conversations.

The Analogy: Imagine trying to learn how to dance by reading a manual on physics. You’d be terrible! To learn to dance, you need to go to a club and watch people move. The researchers didn't give the AI a textbook; they gave it a massive "dance floor" of real human interaction. They even used special recordings so the AI could hear both people in a conversation separately, learning exactly when one person starts and the other stops.

2. Changing the "Brain's Dictionary" (The Tokenizer)

AI doesn't see words; it sees "tokens" (little chunks of data). The original system (called Moshi) was built with an English dictionary. Trying to use an English dictionary to understand Hindi is like trying to play a game of Scrabble using only English tiles to spell out Hindi words—it’s frustrating, slow, and messy.

The Analogy: It’s like trying to cook a traditional Indian meal using a recipe book written in French. You might eventually get it done, but you'll be constantly translating and making mistakes. The team threw out the French recipe book and wrote a brand-new one specifically for Hindi (the Devanagari script), making the AI much faster and smarter.

3. The Two-Step Training (The Learning Process)

They didn't just throw the data at the AI and hope for the best. They used a two-stage approach:

  • Stage 1 (The Deep Dive): They let the AI soak in the massive 26,000-hour library to learn the basic patterns of the language.
  • Stage 2 (The Polishing): They took a smaller, "high-quality" set of conversations (the best, clearest ones) to fine-tune the AI, teaching it to be extra polite and clear.

Does It Actually Work?

The researchers tested the AI to see if it could "keep the ball rolling" in a conversation. Here is what they found:

  • It sounds human: In tests, people rated the AI's naturalness very high. While it's not quite a human yet, it's getting very close.
  • It understands the "Rhythm": They tested if the AI knows when to pause or when to jump in. They found that if they set the "creativity" level just right, the AI mimics human timing—knowing when to wait and when to respond.
  • It’s efficient: Even though the original English model used 7 million hours of data, this Hindi model achieved great results with only 26,000 hours. This proves that quality matters more than quantity.

The Bottom Line

Human-1 is a massive step toward a future where you can talk to your phone in Hindi just like you’re talking to a person on the street—interrupting, laughing, and flowing naturally, without that awkward "robot silence" in between.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →