Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection
This paper critically re-evaluates SemEval-2020 Task 1, the leading benchmark for lexical semantic change detection, by identifying significant limitations in its narrow operationalization of change, substantial data quality issues, and restrictive design choices, ultimately arguing that it should be viewed as a partial tool rather than a definitive measure of progress while calling for future improvements in theory, transparency, and scope.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to build a machine that can read old books and tell you exactly how the meaning of words has changed over the last 200 years. To test if your machine is smart, you need a "driver's license test" for it. In the world of computer science and linguistics, the most famous test for this is called SemEval-2020 Task 1.
For years, researchers have used this test as the gold standard to see who has the best "word-change detector." But in this paper, the authors are acting like a tough car inspector. They are saying: "Hold on a minute. This test isn't just a little flawed; the test itself is driving us in the wrong direction."
Here is a breakdown of their critique using simple analogies:
1. The "Lego Block" Problem (Operationalisation)
The Issue: The test assumes that word meanings are like Lego blocks. You either have a block (a specific meaning), or you don't. If you add a new block or lose an old one, the word has "changed."
The Reality: The authors argue that meaning is more like water flowing in a river. It doesn't jump from one distinct shape to another; it shifts gradually.
- The Metaphor: Imagine the word "nice." In the past, it meant "foolish." Today, it means "pleasant." The test treats this as if the word suddenly swapped one Lego block for another. But in reality, the meaning slowly melted and reformed over centuries.
- The Flaw: By forcing these smooth, flowing changes into rigid "Yes/No" boxes (Did it change? Yes/No), the test misses the subtle, messy, and gradual ways language actually evolves. It's like trying to measure the temperature of a cup of coffee by only asking, "Is it boiling or is it ice?"
2. The "Dirty Glasses" Problem (Data Quality)
The Issue: To test the machine, the organizers gave it a pile of old newspaper scans and books. But these weren't clean scans; they were full of dirt, smudges, and torn pages.
The Reality: The data is full of OCR noise (errors made when computers scan old text).
- The Metaphor: Imagine asking a student to read a history book, but the pages are covered in coffee stains, the ink is faded, and some words are missing entirely.
- Sometimes the computer sees "blockade" but reads it as "bloqiiade."
- Sometimes a sentence gets cut off in the middle, like a radio signal fading out.
- Sometimes the computer mislabels a verb as a noun (calling "landing" a "land").
- The Flaw: If your machine is trying to learn how words change, but the text it's reading is broken, it might think a word changed meaning when it actually just looked weird because of a smudge. It's like blaming the student for failing the test when the teacher handed them a book with half the pages torn out.
3. The "Tiny Sample" Problem (Benchmark Design)
The Issue: The test only asks the machine to check about 30 to 40 specific words per language.
The Reality: This is like judging a chef's entire career based on how well they cook one single egg.
- The Metaphor: If you have a test with 40 questions, getting just one answer wrong changes your score by a huge amount (like 2.5%).
- If the test is too small, the results are just noise. A system might look "better" just because it guessed the right 40 words by luck, not because it's actually smarter.
- It's like a weather forecast that only looks at the temperature in one tiny park in London and claims to predict the weather for the whole world.
- The Flaw: The test is too small to be statistically reliable. It's also limited to only four European languages (English, German, Swedish, Latin), so we can't be sure if the "smart machines" would work on languages like Chinese, Arabic, or Swahili.
The Big Picture: What Should We Do?
The authors aren't saying the test is useless. They are saying it's a rough sketch, not a finished painting.
They suggest we need to:
- Stop thinking in Lego blocks: Accept that meaning changes gradually, like a river, not in jumps.
- Clean the glasses: Make sure the old books are scanned perfectly before we use them to train AI.
- Expand the test: Instead of checking 40 words, check thousands. And don't just check European languages; check the whole world.
In short: The current test is a useful starting point, but if we keep using it as the only measure of success, we might be building AI that is good at passing a broken test, but bad at understanding how human language actually works.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.