← Latest papers
🤖 machine learning

CrediBench: Building Web-Scale Network Datasets for Information Integrity

This paper introduces CrediBench, a web-scale multi-modal dataset comprising eight months of graph, text, and temporal data from over 40 million domains, which significantly enhances the performance of credibility prediction models by capturing structural and dynamic signals often missed by existing claim-level approaches.

Original authors: Emma Kondrup, Sebastian Sabry, Hussein Abdallah, Zachary Yang, Jiaqi Xiong, Kellin Pelrine, James Zhou, Zhijin Guo, Michael M. Bronstein, Jean-François Godbout, Reihaneh Rabbany, Shenyang Huang

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Emma Kondrup, Sebastian Sabry, Hussein Abdallah, Zachary Yang, Jiaqi Xiong, Kellin Pelrine, James Zhou, Zhijin Guo, Michael M. Bronstein, Jean-François Godbout, Reihaneh Rabbany, Shenyang Huang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a massive, bustling city where every website is a building. Some buildings are honest libraries, some are shady back-alley shops, and some are outright traps. Right now, it's getting harder to tell them apart because bad actors are building convincing fake storefronts, and new "buildings" pop up every second.

This paper introduces CrediBench, a giant new tool designed to help us navigate this city and figure out which buildings are trustworthy.

Here is the breakdown of what they did, using simple analogies:

1. The Problem: The Old Maps Are Too Small

Previously, researchers trying to spot fake news or dangerous websites had two main problems:

  • They only looked at the "menu": Most tools just read the text on a single webpage (the menu) to decide if it's a good restaurant. They ignored who the owner was or who they were friends with.
  • The maps were tiny: Existing datasets were like a map of a single neighborhood. They didn't have enough data to see the whole city, especially how the connections between buildings change over time.

2. The Solution: A Giant, Living City Map

The authors built CrediBench, which is like a high-definition, time-lapse map of the entire internet for eight months.

  • The Scale: It's massive. It tracks over 40 million "buildings" (domains) and 1 billion "roads" (hyperlinks) connecting them. To put that in perspective, it's orders of magnitude larger than any previous map of its kind.
  • The Three Layers: Instead of just looking at the text, CrediBench looks at three things at once:
    1. The Text: What is written on the pages? (The menu).
    2. The Structure: Who links to whom? (The neighborhood gossip and alliances).
    3. Time: How do these connections change? (Did a reputable library suddenly start linking to a shady shop last week?).

3. The "Credibility Score"

The researchers taught computers to act like expert inspectors. They created two ways to test the system:

  • The "Trust Meter" (Regression): Instead of just saying "Good" or "Bad," the system gives a score from 0 to 1, indicating how credible a site is.
  • The "Pass/Fail" Test (Classification): A simple "Yes, this is safe" or "No, this is dangerous."

To train these inspectors, they didn't just guess. They gathered a massive list of 662,000 websites that had already been labeled by experts, crowds, and security firms as either "credible" or "unreliable" (covering things like misinformation, malware, and phishing). This is the largest list of its kind ever assembled.

4. The Results: Teamwork Wins

The team ran experiments to see which "sense" was most important: reading the text, looking at the links, or watching the timeline.

  • The Finding: You can't rely on just one sense. A model that only reads text is okay, but a model that only looks at links is also okay.
  • The Winner: The Super-Inspector that combines all three (Text + Links + Time) was the clear champion.
    • On the "Trust Meter" task, it reduced errors by a huge margin (dropping from 16% error down to 10%).
    • On the "Pass/Fail" test, it jumped from being right 56% of the time to being right 85% of the time.

5. Why This Matters

The paper argues that misinformation spreads like a virus through the connections between websites, not just through the words on a page. By looking at the whole city map and how it moves over time, we can spot the "bad neighborhoods" much faster.

In short: The authors built the biggest, most detailed map of the internet's trust network ever created. They proved that to spot a liar, you need to read what they say, see who they hang out with, and watch how their relationships change over time. They made this map and the training tools available for anyone to use, so future researchers can build better "city inspectors" to keep the internet safe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →