What Limits Us? Analyzing Self-Reported Limitations in NLP Research
This paper presents a large-scale analysis of self-reported limitations in ACL and EMNLP papers from 2020 to 2025 using a novel human-AI hybrid coding framework to uncover trends, correlations, and writing patterns in researchers' disclosures.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of artificial intelligence, specifically the branch that teaches computers to understand human language, a new culture of honesty has taken root. For decades, scientists published their breakthroughs with a focus on what worked, often leaving the flaws, failures, and boundaries of their experiments hidden in the fine print or omitted entirely. However, starting in late 2022, the major conferences that host this research made a significant change: they required every accepted paper to include a dedicated section explicitly listing what the work could not do. This mandate transformed the "Limitations" section from an optional footnote into a standard part of the scientific record. The goal was simple but profound: to ensure that anyone reading a study knows exactly where its conclusions might fall short, fostering transparency and preventing others from building on shaky foundations. Now, with thousands of these sections accumulating every year, a vast, unexplored archive of researcher self-reflection has emerged, waiting to be read.
A team of researchers from Chulalongkorn University and Google Research decided to open this archive. Instead of reading the thousands of papers one by one—a task impossible for human hands—they built a new system where human experts and artificial intelligence worked together to read and categorize the text. They analyzed the "Limitations" sections from over 16,000 papers published between 2020 and 2025. Their goal was to understand not just what researchers were saying, but how the nature of these admissions was changing over time, whether different types of research groups admitted to different problems, and if there were hidden patterns in how these limitations were written.
The results revealed a dramatic shift in the landscape of the field. Before the mandatory rule was introduced, fewer than ten percent of papers included a dedicated limitations section. Once the rule took hold, compliance skyrocketed, reaching nearly perfect levels by 2025. However, the content of these sections changed in a surprising way. Before the mandate, researchers most frequently admitted to "methodological constraints," essentially saying their experimental setup had specific technical hurdles. After the mandate, the most common admission became "scope limitation." This is a broader, more general category where authors state that their findings might not apply to larger models, different languages, or real-world scenarios. The researchers noted that this shift happened almost exactly when the policy changed, suggesting that the requirement itself may have encouraged scientists to be more cautious about the general applicability of their work, rather than just listing technical bugs.
Digging deeper into these "scope limitations," the team found that the most rapidly growing concern was the size of the models being tested. As artificial intelligence models have become massive and expensive to run, researchers increasingly admitted that they could only test their ideas on smaller versions of these systems. They wrote that their results might not hold true for the giant, closed-source models used by major tech companies. This reflects a growing divide in the field between those who can afford to train and test on the largest possible systems and those who cannot. Another notable trend was a sharp rise in admissions regarding language coverage. Researchers began explicitly stating that their work was limited to English, acknowledging a long-standing bias in the field where non-English languages are often ignored.
The study also looked at who was writing these limitations. The researchers compared papers from large technology companies with those from universities and smaller institutions. They found that papers from large companies were slightly less likely to admit to "scope limitations" than their counterparts from other institutions. While the difference was small, it suggests that the resources and scale available to big companies might influence how they frame the boundaries of their work. Conversely, researchers from other institutions were more likely to mention data scarcity, perhaps reflecting the challenges of accessing the massive datasets that large companies possess. Despite these differences, the study found a remarkable uniformity across the field; regardless of the specific topic or the author's affiliation, the way limitations were reported remained largely consistent, suggesting a shared culture of caution has taken hold.
Finally, the researchers examined the writing style of these sections, looking for recurring patterns in how authors structured their admissions. They discovered a common rhetorical strategy they called a "soft landing." Often, after stating a limitation, an author would immediately follow it with a suggestion for future work or a justification for why the limitation was unavoidable. This pattern seemed to soften the blow of the admission, turning a negative into a roadmap for what comes next. Another frequent pattern was a "preemptive buffer," where an author would highlight the strengths or successes of their work right before listing its flaws. This technique appeared to cushion the impact of the criticism, ensuring the reader saw the value of the work before being reminded of its boundaries.
The study concludes that while the mandatory policy has successfully forced researchers to be more transparent, it has also changed the nature of that transparency. The field has moved from listing specific technical failures to broadly acknowledging the boundaries of their work. The researchers caution that because these sections are self-reported, they reflect what authors choose to say, not necessarily the full truth of what went wrong. However, by mapping out these trends, the study provides a clear picture of a community in transition, one that is learning to speak more openly about the limits of its own creations. This new habit of explicit self-disclosure offers a rare glimpse into the evolving conscience of the artificial intelligence community, showing a field that is increasingly aware of its own boundaries.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.