← Latest papers
🤖 machine learning

RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations

RiskNet is a large-scale, multilingual dataset of AI risk incidents derived from news sources that employs a structured pipeline for identification, alignment, and multi-dimensional annotation to support empirical research in AI safety, governance, and risk analysis.

Original authors: Leihan Zhang, Wecheng Ye, Xianlong Ma, Haochuan Liu, Yang Li, Qianyu Zhang, Jinliang Chen, Qiang Yan

Published 2026-06-09
📖 5 min read🧠 Deep dive

Original authors: Leihan Zhang, Wecheng Ye, Xianlong Ma, Haochuan Liu, Yang Li, Qianyu Zhang, Jinliang Chen, Qiang Yan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of Artificial Intelligence (AI) as a massive, bustling city. Every day, new AI tools are built and put to work in hospitals, schools, banks, and on our phones. But just like in any big city, things sometimes go wrong. Cars crash, power grids fail, and people get hurt. In the AI city, these "accidents" are called AI risk incidents—things like a chatbot spreading lies, a facial recognition system misidentifying someone, or a self-driving car making a dangerous mistake.

Until now, trying to study these accidents was like trying to map a city's traffic jams by asking people to shout out what they saw from their windows. The information was scattered, messy, in different languages, and often just rumors or opinions rather than facts about specific crashes.

Enter "RiskNet."

Think of RiskNet as a giant, high-tech traffic control center built by researchers from Beijing University of Posts and Telecommunications. It's a massive digital library designed to catch, organize, and understand every AI accident reported in the news.

Here is how it works, broken down into simple steps:

1. The Great Filter (Finding the Accidents)

The researchers started with a mountain of news—hundreds of millions of articles from around the world, in many different languages.

  • The Problem: Most of these articles are just people talking about AI risks (like "AI might be dangerous someday") or general tech news. They aren't about a specific crash that actually happened.
  • The Solution: They used a smart "sieve" (powered by AI itself) to sift through the noise. It asked two questions:
    1. "Is this about AI?"
    2. "Did a specific accident actually happen here?"
    • Analogy: Imagine a bouncer at a club. If you're just talking about how cool the music might be, you don't get in. But if you have a ticket for a specific show that already started, you're let through. RiskNet filters out the "talk" and keeps only the "action."

2. The Detective Work (Connecting the Dots)

Once they found the articles about real accidents, they faced a new problem: One accident often has many stories.

  • The Scenario: Imagine a robot vacuum catches fire in a kitchen. One newspaper in China writes about it. A blog in the US writes about it. A TV station in France reports it. They are all talking about the same fire, but they look like three different stories.
  • The Solution: RiskNet acts like a super-detective. It reads all these different reports and says, "Wait a minute, these three articles are describing the exact same event!" It then glues them together into a single "Incident Cluster."
  • Analogy: Think of it like a puzzle. You have 50 different puzzle pieces from different boxes, all showing parts of the same picture. RiskNet sorts them so you can see the whole picture of the accident, rather than just a confusing pile of fragments.

3. The Labeling System (Organizing the Files)

Now that they have a clean list of unique accidents, they need to organize them so researchers can find patterns.

  • They created a filing cabinet with specific tags for every accident:
    • Who was hurt? (The victim)
    • What went wrong? (The cause)
    • How bad was it? (The severity)
    • Where did it happen? (The location)
    • What kind of AI was involved? (The tool)
  • Analogy: It's like a hospital's emergency room log. Instead of just writing "Patient came in," they write "Patient: John Doe; Injury: Broken leg; Cause: Bicycle accident; Severity: Moderate." This makes it easy to ask questions like, "How many bicycle accidents happened last month?"

What Did They Find?

The result is RiskNet, a dataset containing:

  • Hundreds of millions of news articles scanned.
  • Hundreds of thousands of AI-related risk reports.
  • Over 54,000 unique, verified accident clusters (where multiple news stories were merged into one event).

They discovered that most accidents get very little attention (like a small fender bender), but a few huge, famous accidents get thousands of news stories (like a major city-wide blackout). The data also shows that "lack of capability" (the AI just wasn't good enough) and "cyberattacks" are the most common types of problems.

Why Does This Matter?

The researchers say this dataset is a foundation stone for the future.

  • For Policymakers: It helps them see the real problems, not just the scary headlines, so they can write better rules.
  • For Scientists: It gives them a standardized way to test if new AI systems are safer than old ones.
  • For Everyone: It bridges the gap between "high-level principles" (like "AI should be safe") and the "messy reality" of what actually goes wrong.

Important Note: The paper emphasizes that this is a research tool, not a court of law. It organizes what the news says happened. It doesn't prove legal guilt or assign blame in a courtroom; it just helps us understand the landscape of AI risks so we can study them better.

In short, RiskNet turns a chaotic storm of news reports into a clear, organized map of where AI is stumbling, helping us figure out how to help it walk straight.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →