Do LLMs Know What Luxembourgish Borrows? Probing Lexical Neology in Low-Resource Multilingual Models
This paper introduces LexNeo-Bench, a Luxembourgish lexical neology benchmark, to demonstrate that injecting structured linguistic knowledge graphs into prompts significantly improves large language models' ability to classify lexical borrowings in low-resource contact languages, though detecting neology remains challenging.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, well-read robot that has read almost everything on the internet. You ask it to read a sentence in Luxembourgish, a small language spoken in a tiny country sandwiched between Germany, France, and Belgium. In this country, people often mix words from French, German, and English into their daily speech.
The question the researchers asked was: Does this robot actually understand which words are "native" Luxembourgish and which ones are borrowed from its neighbors?
Here is a simple breakdown of what they found, using some everyday analogies.
The Problem: The Robot is Guessing
The researchers built a test called LexNeo-Bench. Think of it like a multiple-choice quiz where the robot has to look at a specific word in a sentence and decide:
- Is this a native Luxembourgish word?
- Did it come from French?
- Did it come from German?
- Did it come from English?
The Result: When they just asked the robot to do this without any help (like a student taking a test with no notes), it performed barely better than random guessing. Even though the robot had read millions of books, it didn't "know" the rules of how Luxembourgish borrows words. It was like asking someone who has seen a million photos of dogs to identify a specific breed of dog without ever having studied a dog book—they just guessed.
The Solution: Giving the Robot a "Cheat Sheet"
The researchers realized the robot needed a little help. They created a Linguistic Knowledge Graph.
Think of this as a super-detailed cheat sheet or a map. Instead of just giving the robot the sentence, they gave it a specific note card for that exact word that said:
- "This word looks like it comes from French."
- "It follows this specific spelling pattern used in Luxembourgish."
- "Here are three other words that look similar and are definitely French."
The Result: When they gave the robot this cheat sheet, its performance skyrocketed. It went from guessing (about 25–35% correct) to being an expert (71–81% correct).
The Surprising Twist: Smaller Robots Did Better
Usually, in the world of AI, bigger is better. A giant robot with a massive brain (like the 70-billion-parameter model) is expected to beat a smaller one.
But here, the opposite happened.
- The smaller robot (12 billion parameters) became the star student when given the cheat sheet. It followed the instructions perfectly.
- The giant robot (70 billion parameters) actually did worse than the small one when given the same cheat sheet.
Why? The researchers think the giant robot was too "stubborn." It had read so much French and German on the internet that it had a strong internal belief that "If this word looks French, it must be French." When the cheat sheet told it, "Actually, this is a borrowed word adapted into Luxembourgish," the giant robot ignored the note and stuck to its own gut feeling. The smaller robot didn't have those strong pre-existing biases, so it listened to the cheat sheet and got the answer right.
The "New Word" Problem
The researchers also tried to trick the robot into spotting "new words" (neologisms). They asked: "Is this a brand new word, or has it been around for a while?"
The Result: The robot failed miserably at this, even with the cheat sheet.
- The Analogy: The cheat sheet told the robot where a word came from (its origin), but it didn't tell the robot when the word arrived. It's like giving someone a map of a city but not telling them which buildings were built yesterday and which were built 100 years ago. The robot couldn't tell the difference between a "new" borrowed word and an "old" borrowed word.
The Main Takeaway
- AI isn't magic: Just because a robot has read a lot doesn't mean it understands the specific, messy rules of small languages.
- Context is king: If you give the AI a specific, structured guide (like a cheat sheet) about the language's rules, it can learn very quickly.
- Bigger isn't always better: Sometimes, a smaller, more flexible AI listens to new information better than a giant, stubborn one that relies too much on what it already "thinks" it knows.
- Time is hard: AI is good at figuring out where a word comes from, but it's still terrible at figuring out when a word became part of the language.
In short, for small languages like Luxembourgish, you can't just rely on the AI's memory; you have to give it a specific map to navigate the borrowed words.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.