← Latest papers
💬 NLP

Rashid: A Cipher-Based Framework for Exploring In-Context Language Learning

The paper introduces Rashid, a framework that reversibly ciphers high-resource languages to create controllable, resource-rich "unseen" languages, enabling scalable and rigorous evaluation of in-context language learning (ICLL) strategies and overcoming the data and tool limitations typically associated with low-resource languages.

Original authors: Niyati Bafna, Ryan Soh-Eun Shim, Barbara Plank, David Yarowsky, Hale Sirin

Published 2026-03-25
📖 4 min read☕ Coffee break read

Original authors: Niyati Bafna, Ryan Soh-Eun Shim, Barbara Plank, David Yarowsky, Hale Sirin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to teach a super-smart robot how to speak a language it has never heard before, like a rare dialect spoken by a small tribe in the Amazon. The problem? You don't have a dictionary, a grammar book, or any native speakers to help you. The robot is flying blind, and it's impossible to know if it's learning correctly because you can't understand the output either.

This is the big headache researchers face with In-Context Language Learning (ICLL): trying to teach AI new languages without the usual tools.

Enter Rashid, a clever new framework introduced in this paper. The name is inspired by the legendary Caliph Harun al-Rashid, who famously disguised himself as a commoner to walk the streets of Baghdad and understand the people's problems. Similarly, this framework "disguises" a well-known language to trick the AI into thinking it's learning a brand-new one, allowing researchers to study the learning process in a controlled, safe environment.

Here is how it works, broken down with simple analogies:

1. The "Magic Scrambler" (The Cipher)

Instead of trying to find a rare, low-resource language to test on, the researchers took 10 common, high-resource languages (like French, German, Hindi, and Turkish) and ran them through a "Magic Scrambler."

  • The Analogy: Imagine you have a perfect English sentence: "The cat sat on the mat."
  • The Scramble: You apply a secret code where every letter is swapped for another (e.g., 'T' becomes 'X', 'h' becomes 'q'). The sentence becomes: "Xq qat qat on qe mat."
  • The Result: To the AI, this looks like a completely alien language. But to the researchers, it's still French or Hindi underneath. Because they know the secret code, they can instantly "unscramble" the AI's answer to check if it's right.

This is the genius part: They created "fake" unknown languages that actually have perfect dictionaries and grammar books available. This lets them test AI strategies without the usual resource shortages.

2. The Experiment: Teaching the Robot

Once the language is scrambled, the researchers tried to teach the AI to translate it using different "study aids," just like a human student would:

  • The Dictionary: Giving the AI a list of word-for-word translations.
  • The Grammar Book: Giving it rules about sentence structure (like "verbs go at the end").
  • The Examples: Showing it sample sentences of the language being used.

What they found:

  • Reading is easier than writing: The AI got pretty good at understanding the scrambled language (translating it to English) when given a dictionary and grammar rules. It's like reading a foreign menu; you can guess the ingredients.
  • Writing is hard: When asked to generate text in the scrambled language (translating English to the new language), the AI struggled. Even with a dictionary, it couldn't figure out how to put the words together correctly. It was like trying to write a poem in a language you only know the words for, but not the rhythm.

3. The "Pivot" Trick (The Bridge)

The researchers discovered a clever workaround for the writing problem. Instead of going straight from English to the "Scrambled Language," they went through a related language first.

  • The Analogy: Imagine you want to translate a story from English to a rare dialect of Hindi. You don't know the dialect. But you do know standard Hindi.
  • The Strategy: You translate English -> Standard Hindi -> Scrambled Hindi.
  • The Result: Because Standard Hindi and the Scrambled version share the same "skeleton" (grammar and word order), the AI did a much better job. It was like using a bridge to cross a river instead of trying to swim across it in the dark.

4. Why This Matters

This framework is like a flight simulator for language learning.

  • Real Life: Testing AI on real, rare languages is expensive, slow, and risky. If the AI fails, you might not even know why because you don't speak the language.
  • The Rashid Simulator: Researchers can crash the plane (the AI) a thousand times, check the black box (the output), and fix the engine (the strategy) instantly because they can "unscramble" the crash.

The Bottom Line

The paper concludes that while current AI methods are getting better at reading new languages, they are still terrible at speaking them. However, by using this "scrambled language" sandbox, researchers can now rapidly test new ideas—like using related languages as a bridge—to figure out how to finally teach AI to speak the world's thousands of languages fluently.

It's a way to solve the "low-resource" problem by pretending the resources don't exist, so we can learn how to build them for real.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →