Context-Aware LLM Evaluators for Content Moderation Judgments in a Low-Resource Setting
This study demonstrates that while context-aware prompting and retrieval of prior decisions can improve Large Language Model evaluators for German newspaper forum moderation, the choice of model capacity is more critical than the prompting strategy, with the largest model outperforming supervised systems on rare removal categories without requiring annotated training data.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the bustling digital town squares of the internet, where millions of comments are posted every day, a quiet crisis of scale has emerged. Human moderators, the trained professionals who once read every post to decide what stays and what goes, are overwhelmed. The sheer volume of content has made it impossible for people to keep up, yet the rules for what constitutes a harmful or off-topic post are often subtle, depending heavily on the specific conversation and the community's norms. To solve this, researchers have turned to large language models, a type of artificial intelligence capable of understanding and generating human-like text. These models are increasingly being asked to act as judges, assigning labels to posts and making moderation decisions that were once the exclusive domain of humans. However, a critical question remains: can a machine truly understand the nuance of a heated online argument, or does it need to see the whole picture to make a fair call?
This question drove a team of researchers from Austrian universities to investigate how well these artificial judges perform when they are given different amounts of information. They focused on a specific, challenging environment: the comment sections of a German-language newspaper. In this setting, deciding whether a comment is off-topic, discriminatory, or simply a personal story requires more than just reading the words on the screen; it requires understanding the context of the article it replies to and the history of the conversation it is part of. The researchers wanted to know if giving the artificial intelligence more context—like the full news article or the entire chain of replies—would help it mimic the decisions of human experts. They also tested a different approach: instead of feeding the model more text to read, could they show it examples of how humans had judged similar comments in the past?
To find the answers, the team used a massive collection of over one million posts from the Austrian newspaper Der Standard, which included a smaller, carefully annotated set of nearly 12,000 comments that had been graded by professional human moderators. They tested three different artificial intelligence models of varying sizes, ranging from a smaller, faster model to a much larger, more complex one. For each model, they ran a series of experiments where they changed the information provided in the prompt. In some cases, the model saw only the single comment in question. In others, it saw the comment plus the headline of the article, the full text of the article, or the entire thread of previous replies. They also tested a method where the model was shown two short, retrieved examples of how humans had judged similar comments before, one that was kept online and one that was removed.
The results revealed a nuanced picture of how these digital judges think. The researchers found that context does help, but not in the way one might expect. Simply dumping more text into the system does not guarantee better results. In fact, for some types of judgments, such as spotting discriminatory language or inappropriate insults, adding the full text of the news article actually confused the model and made its performance worse. The model did better when it saw the conversation history—the replies and the flow of the discussion—because that is where the true meaning of a comment often lies. However, for other categories, like deciding if a comment is off-topic, seeing the original article was essential. There was no single "best" amount of information that worked for every situation; the right amount of context depended entirely on what the model was being asked to judge.
Perhaps the most surprising discovery was that showing the model past human decisions was just as effective as showing it the full article and conversation history, but with a major advantage: it was much shorter and faster. When the researchers gave the model two short examples of how humans had judged similar comments in the past, the model performed nearly as well as when it was given the entire news article and the full conversation thread. This retrieval method was also more reliable; it improved the model's performance across all categories, whereas adding more text sometimes hurt performance. It turned out that seeing how a human had solved a similar problem in the past was a powerful shortcut that allowed the model to learn the community's standards without needing to read pages of background text.
The size of the artificial intelligence model proved to be the most critical factor of all. The largest model, with 31 billion parameters, was the only one that could consistently match the judgment threshold of the human experts, making balanced decisions that were neither too strict nor too lenient. The smaller models struggled significantly. The smallest model, with just 1 billion parameters, often failed to improve when given more context or examples, and sometimes performed worse with them. The mid-sized model could be made to flag almost every comment as problematic, catching nearly everything but also making many false alarms. This suggests that while smaller models might be useful as a first filter to catch obvious issues, they are not yet ready to replace human judgment on their own. Only the largest models demonstrated the necessary capability to understand the subtle boundaries of community norms.
When compared to previous computer systems that were trained specifically on this data, the new approach showed a distinct strength in the areas that matter most for safety. The artificial judges were significantly better at identifying the rare, harmful posts that lead to removal, such as discriminatory or off-topic content, outperforming older systems that had been trained on thousands of examples. However, the older, trained systems were still better at identifying constructive content, like personal stories or well-argued points. This difference highlights a key reality: the new approach works best where human training data is scarce and expensive to produce. For the rare, difficult cases that define the safety of a community, the artificial judge, guided by a few smart examples, can do a better job than a system trained on limited data. But for common, positive interactions, the old methods still hold an edge.
Ultimately, the study suggests that the future of content moderation is not about replacing humans with a single, perfect machine, but about finding the right tools for the right job. The research indicates that for newsrooms and online communities, the best strategy involves using the largest available models and relying on a method that shows them how humans have judged similar situations in the past, rather than overwhelming them with endless text. This approach allows for scalable, reliable moderation even in languages and communities where there is no massive database of labeled examples to train on. It offers a path forward where technology can handle the volume of the internet while respecting the subtle, context-dependent nature of human conversation, provided the right model is chosen and the right information is given.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.