← Latest papers
📄 medicine

Evaluating Agreement Between Human Consensus and Large Language Model Ratings of Suicide Reporting

This study demonstrates that large language models can effectively rate suicide news articles using the TEMPOS framework with agreement comparable to human raters, suggesting AI is a viable tool for scaling the monitoring of responsible suicide reporting when paired with human oversight.

Original authors: Andrew Sarkin, Bilal Malik, Elizabeth Doery

Published 2026-08-12
📖 4 min read☕ Coffee break read

Original authors: Andrew Sarkin, Bilal Malik, Elizabeth Doery

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where the news acts like a giant, invisible mirror. Sometimes, this mirror reflects stories in a way that helps people find hope and solutions; other times, it can accidentally show a distorted image that makes a bad situation feel even worse. Scientists call this the "Werther Effect," where certain ways of reporting on suicide can sadly lead to more people hurting themselves, and the "Papageno Effect," where careful, hopeful stories can actually save lives. Because of this, health experts have created a set of "rules of the road" for journalists to follow, like avoiding sensational headlines or including help lines. But here's the tricky part: checking if millions of news articles follow these rules is like trying to count every single grain of sand on a beach by hand. It takes forever, it's exhausting, and reading so many sad stories can be emotionally heavy for the people doing the counting. This is where the new idea of using "AI judges" comes in. Could a super-smart computer program, trained to read these rules, do the counting for us? That's the big question this study set out to answer.

The researchers, a team from the University of California, San Diego, decided to put this idea to the test using a specific checklist called TEMPOS (Tool for Evaluating Media Portrayals of Suicide). Think of TEMPOS as a 10-point scorecard where a news article gets points for being safe and respectful, and loses points for being sensational or dangerous. To see if computers could play this game as well as humans, they gathered 42 real news stories about suicide deaths. They then lined up two teams of judges: a team of seven trained human experts and a team of five different AI models (the "brains" behind tools like ChatGPT, Grok, and others).

The humans and the AIs read the same stories and gave them scores independently, without talking to each other. First, the researchers looked at how well the humans agreed with each other. They found that while one single human judge might see things a little differently than another, when you averaged out all seven humans, they created a very stable "Gold Standard" score (which they named HumanGold). This is like having a panel of seven food critics; one might love the salt, another might hate it, but their average opinion is usually a very reliable taste of the dish.

Next, they asked: "How close did the AI judges get to the HumanGold?" The results were surprisingly good. The individual AI models showed a "moderate to strong" agreement with the human team. One model, called Grok, was particularly sharp, matching the human consensus almost as well as the best human judges did. But the real magic trick happened when the researchers combined the top three AI models into a single "super-judge" team, which they dubbed AISilver. This combined AI team matched the human consensus with a correlation of 0.797, which is almost as good as the best single AI model.

However, the study also found some important "watch out" signs. The AI (and even the humans) were much better at spotting concrete, easy-to-see things, like "Did the article include a phone number for help?" or "Did it describe the method of suicide?" These are like spotting a red car in a parking lot; it's either there or it isn't. But the AI struggled more with the "squishy," interpretive parts of the scorecard, like "Did the article make suicide look glamorous?" or "Was the language sensational?" These are like trying to judge the mood of a song; it's much harder for a computer (and even humans) to agree perfectly on these feelings. Also, the researchers noticed that the AI models sometimes gave scores that were a bit too high on average, like a teacher who is just a little too generous with A's.

So, what's the final verdict? The paper suggests that AI isn't a perfect replacement for humans yet, but it is a very promising "assistant." It can't make the final, definitive judgment on its own, especially on the tricky, emotional parts of the story. But if you use a team of AI models (like the AISilver team) to do the heavy lifting of scanning thousands of articles, and then have a small team of humans check the results or handle the tricky cases, it could be a game-changer. This approach could help monitor news coverage on a massive scale without burning out human experts, ensuring that the news mirror reflects stories in a way that protects and helps rather than harms. The study concludes that while we shouldn't trust the AI blindly, pairing it with human oversight could be the key to keeping suicide reporting safe and responsible in the digital age.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →