← Latest papers
🤖 AI

InfiCoEvalChain: A Blockchain-Based Decentralized Framework for Collaborative LLM Evaluation

To address the statistical instability and opacity of centralized LLM benchmarking, this paper proposes **InfiCoEvalChain**, a blockchain-based decentralized framework that leverages heterogeneous compute nodes and a reward-driven consensus mechanism to provide more stable, reliable, and transparent model evaluations.

Original authors: Yifan Yang, Jinjia Li, Kunxi Li, Puhao Zheng, Yuanyi Wang, Zheyan Qu, Yang Yu, Jianmin Wu, Ming Li, Hongxia Yang

Published 2026-02-10
📖 4 min read☕ Coffee break read

Original authors: Yifan Yang, Jinjia Li, Kunxi Li, Puhao Zheng, Yuanyi Wang, Zheyan Qu, Yang Yu, Jianmin Wu, Ming Li, Hongxia Yang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Problem: The "Fickle Exam" Problem

Imagine you are a student preparing for a massive final exam. You take a practice test on Monday and get a 95%. You take the exact same test on Tuesday, but because the room was slightly warmer, the lighting was different, and you were using a different brand of pen, you get an 85%.

In the world of Artificial Intelligence (AI), this is actually happening. Currently, when scientists want to see how "smart" an AI model is, they use "benchmarks" (standardized tests). But because AI generates answers using a bit of randomness (like rolling dice), and because different computers have different "temperatures" or hardware, the scores jump around wildly.

Right now, the gap between the "smartest" AI and the "second smartest" AI is often smaller than the error margin of the test itself. It’s like trying to rank the fastest runners in the world, but the stopwatch is broken and changes its reading every time you press start. We don't actually know who is winning; we just know who got a "lucky" reading.


The Solution: InfiCoEvalChain (The "Global Jury")

The researchers proposed a new system called InfiCoEvalChain. Instead of one big company (like Google or OpenAI) running the test in their own private lab (a "Black Box"), they want to turn the test into a Global, Decentralized Jury.

Here is how it works using three metaphors:

1. The Diverse Jury (Hardware & Parameter Diversity)

Instead of one professor grading a paper, imagine a jury of 1,000 people from all over the world. Some are sitting in snowy mountains, some are in tropical jungles, some are using high-end supercomputers, and some are using old laptops.
By having everyone grade the same AI at the same time, the "noise" (the weirdness caused by the environment) cancels itself out. If the AI is truly smart, it will perform well for everyone. If it just got "lucky" because of a specific computer setting, that luck won't hold up across the whole jury.

2. The Sealed Envelope (The Blockchain & Commit-Reveal)

To make sure no one cheats, they use Blockchain technology. Think of this like a "Sealed Envelope" system.

  • Step 1 (Commit): A juror grades the AI and puts their score in a sealed, locked envelope. They tell the world, "I have a score, and here is a digital fingerprint to prove it's locked."
  • Step 2 (Reveal): Only after everyone has submitted their envelopes does the jury open them all at once.
    This prevents "copycatting," where one juror waits to see what the smartest person said and then just copies them to get a reward.

3. The "Fair Play" Reward (The Incentive Mechanism)

How do you keep people from lying or being lazy? You use a Smart Reward System.
Imagine a game show where you get paid based on how close your answer is to the "consensus" (the middle ground of the group). If you provide a score that is wildly different from everyone else (an outlier), the system assumes you are either a cheater or your computer is broken, and you get almost nothing. If you are part of the honest majority, you get a fair share of the prize.


The Result: From "Blurry" to "High-Definition"

The researchers tested this, and the results were massive.

Before, the scores were "blurry"—they fluctuated so much you couldn't trust them. After using this decentralized "Jury" method, the fluctuations dropped significantly (in one test, the error margin dropped from 1.67 down to 0.28).

In short: They have moved AI evaluation from a "shaky, private measurement" to a "rock-solid, public consensus." It’s like moving from a blurry, handheld camera photo to a crystal-clear, professional photograph. Now, when a leaderboard says "Model A is better than Model B," we can actually believe it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →