← Latest papers
💬 NLP

Far Out: Evaluating Language Models on Slang in Australian and Indian English

This paper evaluates the slang comprehension capabilities of seven state-of-the-art language models on Indian and Australian English using web-sourced and synthetic datasets, revealing significant performance disparities between discriminative and generative tasks and highlighting that models generally perform better on Indian English than Australian English.

Original authors: Deniz Kaya Dilsiz, Dipankar Srirag, Aditya Joshi

Published 2026-02-19
📖 4 min read☕ Coffee break read

Original authors: Deniz Kaya Dilsiz, Dipankar Srirag, Aditya Joshi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot librarian who has read almost everything on the internet. You ask this robot to tell you a joke, finish a story, or guess what word comes next in a sentence. Usually, it's amazing. But what happens when you ask it to understand slang?

This paper is like a "slang test" for seven of the smartest AI robots currently available. The researchers wanted to see if these robots could understand the unique, colorful, and often confusing slang used by people in Australia and India.

Here is the breakdown of their experiment, explained simply:

1. The Setup: Two Different "Slang Dictionaries"

The researchers didn't just guess what slang to test; they built two specific "test banks" to see how the robots performed:

  • The "Real World" Bank (WEB): They went to Urban Dictionary (a website where real people define slang) and grabbed 377 real examples of Australian and Indian slang. Think of this as testing the robot with real street interviews.
    • Example: "Far out" (Australian for "Wow, that's crazy!") or "Prepone" (Indian for "to move a meeting to an earlier time").
  • The "Made-Up" Bank (GEN): They asked a super-advanced AI to invent 1,492 new scenarios where these slang words would be used. Think of this as testing the robot with fictional stories written by a writer.

2. The Three Tests

The researchers gave the robots three different types of challenges:

  • Test A: The "Fill-in-the-Blank" (Target Word Prediction):

    • The Task: "I can't believe how much that concert ticket cost! It's absolutely ______!" (The robot has to guess the missing word).
    • The Catch: The robot has to pull the word out of thin air from its entire memory.
    • The Result: Disaster. The robots were terrible at this. They guessed random words or standard English instead of the slang. It's like asking someone to improvise a rap song and them just saying, "Uh, hello?"
  • Test B: The "Guided Fill-in-the-Blank" (TWP):*

    • The Task: Same as above, but the robot gets a hint: "Remember, this is Australian slang."
    • The Result: Still pretty bad. The hint didn't help much.
  • Test C: The "Multiple Choice" (Target Word Selection):

    • The Task: "I can't believe how much that concert ticket cost! It's absolutely ______!"
      • A) Expensive
      • B) Far out
      • C) Delicious
      • D) Blue
    • The Result: Much better! When the robot just had to pick the right answer from a list, it got it right almost half the time.

3. The Big Surprises

The study found four main things that tell us a lot about how AI thinks:

  • Recognition is easier than Creation: The robots are like a person who can recognize a song when they hear it on the radio but can't sing a single note of it themselves. They are much better at choosing the right slang than making it up.
  • Real vs. Fake: The robots did slightly better on the "Real World" (WEB) examples than the "Made-Up" (GEN) ones. This suggests the robots might have "cheated" by memorizing the real examples from their training data, rather than truly understanding the concept.
  • India vs. Australia: The robots were generally better at understanding Indian English slang than Australian slang. This might be because there is just more data about Indian English on the internet for the robots to learn from.
  • The "Literal" Trap: When the robots failed, they didn't usually make up nonsense. Instead, they gave boring, literal answers.
    • If the slang was "Far out" (meaning amazing), the robot guessed "Great."
    • If the slang was "Bogan" (a specific Australian term for a rough person), the robot guessed "Heavy metal band."
    • They understood the meaning but missed the flavor and the culture.

4. Why Does This Matter?

Imagine you are building a customer service bot for a bank in Mumbai or a content filter for a social media app in Sydney. If your AI doesn't understand local slang, it might:

  • Miss a joke and think a user is being rude.
  • Fail to understand a complaint written in casual language.
  • Make people feel like the technology doesn't "get" them.

The Bottom Line

This paper is a reality check. Even though AI is getting smarter, it still struggles with the messy, fun, and cultural parts of human language. It's great at reading the dictionary, but it's still learning how to speak the "street."

The takeaway: If you want an AI to understand slang, don't ask it to write a poem; give it a multiple-choice quiz. And even then, it might still need a human to explain the joke.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →