← Latest papers
💬 NLP

Making Large Language Models Speak Tulu: Structured Prompting for an Extremely Low-Resource Language

This paper demonstrates that structured prompting techniques, including explicit grammar documentation, negative constraints, and synthetic data generation, can enable large language models to achieve high grammatical accuracy in Tulu, an extremely low-resource language, without requiring model fine-tuning.

Original authors: Prathamesh Devadiga, Paras Chopra

Published 2026-02-18
📖 5 min read🧠 Deep dive

Original authors: Prathamesh Devadiga, Paras Chopra

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, super-smart robot (a Large Language Model) that has read almost every book, website, and article in the world. But there's a catch: it has only read about 10% of the world's languages. The other 90%? It's basically a blank slate.

Now, imagine you want this robot to speak Tulu, a beautiful language spoken by 2 million people in India, but which barely exists online. If you just ask the robot, "Please speak Tulu," it will panic. Since it knows Tulu is similar to Kannada (a language it knows very well), it will accidentally start speaking Kannada instead. It's like asking someone to speak French, but because they know Spanish so well, they accidentally start speaking Spanish with a French accent.

This paper is a story about how the authors taught this robot to speak Tulu without teaching it new words or retraining its brain. Instead, they used a very specific, structured "instruction manual" (a prompt) to guide it.

Here is how they did it, explained with some everyday analogies:

1. The Problem: The "Imposter" Language

The robot is like a student who studied hard for a Spanish exam but was asked to take a Portuguese test. Because the two languages look and sound similar, the student keeps writing Spanish words, thinking they are Portuguese.

  • The Issue: When the robot sees Tulu written in the Kannada script (which looks like Kannada), it assumes, "Oh, this is Kannada!" and starts using Kannada words. The authors call this "Vocabulary Contamination."

2. The Solution: The "Super-Prompt"

Instead of feeding the robot thousands of new books (which don't exist for Tulu), the authors built a massive, detailed instruction manual to give the robot before it answers. Think of this manual as a GPS navigation system for the robot's brain.

They used four main tools to build this GPS:

A. The "Accent Switch" (Romanization)

The Tulu language is usually written in the Kannada script, which confuses the robot. The authors decided to write Tulu using the English alphabet (Romanization), but with a special twist.

  • The Analogy: Imagine the robot is a chef who only recognizes ingredients in French. If you give them a list in French, they might grab the wrong spice. But if you give them a list in English with specific notes like "Use this specific type of salt, not the regular one," they get it right.
  • The Result: By writing Tulu in a special English-style alphabet, the robot realized, "Wait, this isn't Kannada! This is a different language!" This reduced the robot's confusion by a huge margin.

B. The "Do Not Enter" Signs (Negative Constraints)

This was the most powerful tool. The authors didn't just say, "Speak Tulu." They said, "DO NOT use these 50 specific Kannada words. If you want to say 'I', do NOT use 'naanu' (Kannada); you MUST use 'yān' (Tulu)."

  • The Analogy: It's like telling a child, "Don't eat the red candy." If you just say "Eat the green candy," the child might still grab the red one because it's right there. But if you explicitly put a "STOP" sign on the red candy, the child avoids it.
  • The Result: This "negative constraint" stopped the robot from using the wrong words 12–18% more often than any other method.

C. The "Cheat Sheet" (Grammar Documentation)

Since the robot has never seen Tulu grammar before, the authors wrote a mini-textbook inside the prompt. They explained how verbs change, how to say "I," "you," and "we," and the order of words in a sentence.

  • The Analogy: It's like giving the robot a "Cheat Sheet" for a test it didn't study for.
  • The Result: This helped the robot construct sentences correctly, though it worked better for some robot models than others.

D. The "Self-Check" (Self-Verification)

Finally, they told the robot: "Before you hit 'Send,' double-check your work. Did you use the forbidden words? Is the grammar right?"

  • The Analogy: This is like a student reviewing their essay before handing it in to the teacher.
  • The Result: This final step cleaned up the remaining mistakes, boosting the accuracy to 85%.

3. Did it Work?

The results were impressive.

  • Before: The robot was speaking 80% Kannada (the wrong language) and only 18% correctly.
  • After: The robot was speaking 85% correctly Tulu, with only 5% "contamination" from Kannada.

They even did a "lie detector test" (called a falsification experiment). They gave the robot fake, wrong grammar rules and asked it to follow them. The robot followed the fake rules and got terrible results. This proved that the robot wasn't just guessing; it was actually reading and using the instructions they gave it.

4. The Catch (Limitations)

While the robot can now hold a conversation, it's not perfect.

  • The "Robot Voice": The sentences are grammatically correct, but they sound a bit stiff, like a translation rather than a natural conversation.
  • The Script Issue: The authors used English letters to write Tulu because it worked best for the computer. However, real Tulu speakers use Kannada or Tigalari scripts. The paper warns that we shouldn't force English letters on a culture just because it's easier for the AI.
  • Not for Serious Stuff: Because the robot still makes about 15% of mistakes, you shouldn't use this for legal advice, medical diagnosis, or teaching kids. It's great for chatting or exploring, but not for high-stakes jobs.

The Big Takeaway

This paper shows that you don't always need to build a new brain (retrain a model) to teach a robot a new language. Sometimes, you just need to give it a really good, very specific set of instructions (a prompt) that tells it exactly what to do and, more importantly, what NOT to do.

It's a clever, low-cost way to help preserve and speak languages that the digital world has mostly forgotten.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →