NLG Evaluation: Past, Present, Future
This paper traces the evolution of Natural Language Generation (NLG) evaluation from its linguistics-focused origins in 1990 to its current machine learning-driven experimental paradigm, while predicting that future assessment will increasingly prioritize impact, qualitative, and safety metrics as NLG technology becomes ubiquitous.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine Natural Language Generation (NLG) as a chef who writes recipes and cooks meals using computers instead of stoves. For the last 30 years, the way we judge this chef's work has changed completely, and it's about to change again. This paper, written by Ehud Reiter, is a timeline of how we've learned to taste-test these digital meals.
Here is the story of NLG evaluation, broken down into three acts: The Past, The Present, and The Future.
Act 1: The Past (1990 – 2010)
The "Chef's Note" Era
- 1990: The Art Critic Approach
Back in 1990, when the author got his PhD, nobody really "tested" the cooking with numbers. If a chef claimed their new recipe was better, they didn't serve it to 100 people to see if they liked it. Instead, they wrote a long essay explaining why the recipe was good, using fancy linguistic or engineering arguments. It was like a food critic saying, "This sauce is beautiful because the grammar of the ingredients flows well," without actually asking anyone if it tasted good. - 2000: The Taste-Test Explosion
By 2000, people realized they needed to actually taste the food. Researchers started trying all sorts of ways to judge the meals: asking humans to rate them, checking if the food solved a specific problem, or using early computer scores. It was a bit chaotic—everyone was trying different methods, and there was no single "gold standard" yet. - 2010: The Standardized Menu
By 2010, the kitchen got organized. The community started using "Shared Tasks," which were like cooking competitions where everyone had to make the same dish. They agreed on specific rulers to measure the food:- The Ruler (Metrics): They used computer scores like BLEU and ROUGE. Think of these as comparing the new dish to a "perfect reference dish" and counting how many words matched.
- The Panel (Human Ratings): They also asked humans to rank dishes. This became the standard way to judge if the food was actually tasty, because the computer rulers sometimes got it wrong.
Act 2: The Present (2026)
The "Super-Chef" Era
Fast forward to 2026. The chef (now powered by massive AI models called LLMs) has become incredibly good. The food is so perfect that the old ways of judging it are breaking down.
- The Problem with the Old Ruler:
In the past, we compared the AI's food to a human-written "reference dish." But now, the AI's food is often better than the human reference. You can't judge a masterpiece by comparing it to a sketch; the sketch just looks bad by comparison. - The New Judge (LLM-as-Judge):
Since the food is so complex, we started using one AI to judge another AI. It's like hiring a super-taster who is also a robot to critique the meal. This works well sometimes, but it's risky if the robot taster is biased or confused. - The Expert Panel:
Regular people (crowdworkers) aren't good enough to spot subtle errors in these high-quality meals anymore. Sometimes they even cheat by using AI to do the judging! So, we now need experts (like doctors or lawyers) to look closely at the food for specific problems, like "Did the chef hallucinate a fake ingredient?" or "Is this advice safe?" - The Safety Check:
Because these AI chefs are now working in hospitals and law firms, we can't just care about how "average" the meal is. We need to check for worst-case scenarios. If a medical AI gives perfect advice 99.9% of the time but gives deadly advice 0.1% of the time, that's a disaster. We need to hunt for those rare, dangerous mistakes.
Act 3: The Future (2036)
The "Real-World Impact" Era
Looking ahead to 2036, the author argues that we need to stop just measuring how well the chef cooks in a test kitchen and start measuring how the food affects the real world.
- Impact Evaluation (The Health Check):
Right now, we mostly ask, "Did the AI get the right answer?" In the future, we need to ask, "Did using this AI actually help the patient?" or "Did it save the company money?" This is like moving from judging a recipe in a lab to seeing if the restaurant actually keeps customers happy and healthy in the real world. - Qualitative Evaluation (The Storytelling):
We will need to stop relying only on numbers (statistics) and start listening to stories. We'll use interviews and focus groups to understand why people liked or hated the experience. Numbers tell us what happened; stories tell us how it felt. - Safety Evaluation (The Guardrails):
Safety will become the most important thing. Governments might start stepping in to set rules, just like they do for cars or medicine. We will need to constantly monitor the AI to make sure it doesn't go rogue, especially when hackers try to trick it or when it faces weird, unpredictable situations.
The Big Takeaway
The paper concludes that the field is moving from checking boxes (did the numbers match?) to checking the real world (did this actually help or hurt people?).
The author warns that the current research culture is a bit lazy; everyone wants to publish quick results, and reviewers often ignore whether the tests are actually good. To move forward, the community needs to be more rigorous, stop cheating, and learn how to test these powerful tools in the messy, complex reality where they will actually be used.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.