← Latest papers
💬 NLP

OasisSimp: An Open-source Asian-English Sentence Simplification Dataset

This paper introduces OasisSimp, a new open-source multilingual sentence simplification dataset covering five Asian languages (including three with no prior datasets), which serves as both a valuable resource and a challenging benchmark to reveal the performance disparities and limitations of current large language models in low-resource settings.

Original authors: Hannah Liu, Muxin Tian, Iqra Ali, Haonan Gao, Qiaoyiwen Wu, Blair Yang, Uthayasanker Thayasivam, En-Shiun Annie Lee, Pakawat Nakwijit, Surangika Ranathunga, Ravi Shekhar

Published 2026-03-17
📖 5 min read🧠 Deep dive

Original authors: Hannah Liu, Muxin Tian, Iqra Ali, Haonan Gao, Qiaoyiwen Wu, Blair Yang, Uthayasanker Thayasivam, En-Shiun Annie Lee, Pakawat Nakwijit, Surangika Ranathunga, Ravi Shekhar

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to explain a complex recipe to a child. You wouldn't use words like "sauté," "deglaze," or "emulsify." Instead, you'd say, "fry the onions until they're soft," "add the wine to scrape up the tasty bits," and "mix it all together until it's smooth."

This is the heart of Sentence Simplification: taking complicated, grown-up text and rewriting it so everyone can understand it, without losing the original meaning.

For a long time, scientists have been very good at doing this for English. They have huge libraries of "hard sentences" and their "easy versions" to teach computers how to do the job. But for many other languages, especially those spoken in Asia and the Middle East, this library was empty. It was like trying to teach a chef to cook Italian food when you only have a cookbook for French cuisine.

This paper introduces OasisSimp, a new project designed to build that missing library. Here is a simple breakdown of what they did and what they found:

1. The Problem: The "Language Desert"

Think of the world's languages as different gardens.

  • English is a lush, well-watered garden with plenty of tools and maps (data).
  • Thai, Pashto, Tamil, and Sinhala are like beautiful gardens that have been left in the desert. They have rich soil and potential, but no one has built the irrigation systems (datasets) to help plants grow there.

Because of this, computers (AI) struggle to simplify text in these languages. They either get confused or make up nonsense.

2. The Solution: Building the "Oasis"

The researchers decided to build an Oasis (hence the name OasisSimp) in the middle of this desert. They created a new dataset containing pairs of sentences in five languages: English, Sinhala, Tamil, Thai, and Pashto.

  • How they did it: Instead of letting a computer guess, they hired real humans who are native speakers of these languages.
  • The Process: Imagine a team of editors. They took complex sentences from news articles, government documents, and Wikipedia. Then, they manually rewrote them into simpler versions, following strict rules:
    • Cut the fluff: Remove unnecessary details.
    • Swap the words: Replace big, scary words with simple ones.
    • Break it up: Turn one long, confusing sentence into two or three short, clear ones.
  • The Result: They created a "Gold Standard" library. Now, for the first time, researchers have a reliable map to teach computers how to simplify text in these specific languages.

3. The Test: Putting the AI to Work

Once they built the library, they wanted to see if current AI models (the "smart robots" we hear about) could actually use it. They picked 8 different AI models and gave them a test: "Here is a hard sentence; please make it easy."

They tested the AI in two ways:

  • Zero-Shot: The AI had to guess how to do it with no examples (like being handed a new tool with no instructions).
  • Few-Shot: The AI was shown a few examples first (like being shown how to use the tool once or twice before trying).

4. The Results: A Mixed Bag

The results were like a race where some runners had running shoes and others had to run in sand.

  • The English Runner: The AI models were excellent at simplifying English. They knew the rules because they had seen millions of English examples before.
  • The Low-Resource Runners: When the AI tried to simplify Pashto, Thai, or Tamil, it stumbled.
    • The "Zero-Shot" Struggle: Without examples, the AI often failed. It would delete important facts, keep the hard words, or just output gibberish. It was like a tourist trying to order food in a language they don't speak without a phrasebook.
    • The "Few-Shot" Boost: However, when the researchers gave the AI just one or five examples of how to simplify a sentence, the AI got much better. It was like giving that tourist a phrasebook; suddenly, they could communicate!
    • The Star Performer: One model, called Gemma, seemed to be the most adaptable. It handled the different languages better than the others, acting like a polyglot who could pick up new languages quickly.

5. Why This Matters

This paper is a big step forward for digital equality.

  • For People: It means that in the future, people with reading difficulties, children, or non-native speakers in countries like Sri Lanka, Thailand, or Afghanistan will be able to access complex news, legal documents, and medical advice in their own language, written simply.
  • For Science: It highlights that while AI is powerful, it still needs help. It can't just "guess" how to simplify text in a language it hasn't seen much of. It needs human-made examples to learn the nuances.

The Bottom Line

The authors built a new Oasis of data to help computers learn how to speak simply in languages that were previously ignored. They found that while current AI is smart, it needs a little nudge (a few examples) to do a good job in these languages. This work paves the way for a future where information is accessible to everyone, no matter what language they speak or how complex the original text might be.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →