← Latest papers
💻 computer science

StableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction

StableEval Arena is a cost-aware benchmark framework that evaluates LLM-backed agentic systems on stablecoin peg-risk prediction by measuring a joint spectrum of forecast quality, operational reliability, and computational cost, revealing that while agents reliably produce structured outputs, they struggle to accurately predict rare severe-stress and sustained-depeg events.

Original authors: Sean Wan, Dongping Liu, Luyao Zhang

Published 2026-08-07
📖 4 min read☕ Coffee break read

Original authors: Sean Wan, Dongping Liu, Luyao Zhang

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the captain of a massive spaceship fleet, and your cargo holds are filled with "digital gold" that is supposed to always be worth exactly one dollar. This isn't just any gold; it's a special kind of money called a stablecoin, designed to stay perfectly steady so people can buy coffee, pay rent, or trade assets without worrying that their money might suddenly vanish or double in value overnight. But here's the catch: even though these coins are supposed to be boringly stable, they can sometimes panic. When the market gets scary, the price can wobble, drop below a dollar, or even crash completely. This is called "de-pegging," and if it happens, it's like your spaceship's fuel gauge suddenly lying to you.

To keep the fleet safe, we need a lookout who can spot trouble before it happens. In the past, we used simple math or human experts for this. But now, we have Agentic AI—super-smart computer programs that can think, reason, and make decisions on their own, kind of like a digital co-pilot. The big question is: Can these AI pilots actually do a better job than a simple rulebook at spotting a storm before it hits? We need to know if they are truly trustworthy, not just if they can follow instructions. If an AI says, "All clear!" right before the ship crashes, it doesn't matter how fast it is or how pretty its report looks; it's a failure.

This is exactly what the paper "StableEval Arena" sets out to test. The researchers built a giant, cost-aware training ground (an "arena") to see if these AI agents can predict when a stablecoin is about to lose its dollar-peg. They didn't just ask the AI to guess the future price; they asked it to act like a risk manager. The AI had to look at the last 30 days of price history and market chatter, then predict what would happen in the next seven days. Would the coin stay Stable? Would it enter a Watch state (a little shaky)? Or would it go into Stress (a full-blown crisis)?

The researchers ran two different kinds of tests. First, they created a "stress-test" version with 120 cases where trouble was more common, just to see if the AI could spot the danger signs when they were right in front of it. Then, they ran a "real-world" version with 507 cases that looked exactly like the natural world, where stablecoins are usually calm and only rarely have problems. They tested six different AI agents against some very simple baselines (like a rule that just says "assume everything is fine" or a basic math model).

The results were a mix of good news and a serious reality check. On the good side, the AI agents were incredibly reliable at following the rules. They produced perfectly formatted reports, didn't crash, and did it all for a very low cost. They were great at saying, "Hey, things look a little wobbly today." However, when it came to the most important job—spotting a rare, catastrophic crash—the AI agents mostly missed the mark. In the big test of 507 cases, the AI missed 19 out of 21 actual stress events. They were so good at following the protocol that they produced valid reports, but those reports often said "All clear" right before the coin actually crashed.

In fact, in the real-world test, even a simple computer program that just looked at past trends struggled to predict the rare crashes, missing the majority of them just like the AI. The fancy AI agents did not dominate or outperform these classical methods; in fact, none of the agents could reliably detect the rare, high-stakes stress events that mattered most. The paper concludes that while these AI agents are excellent at being obedient and efficient, they aren't yet trustworthy enough to be the sole guardians of our digital money. They can follow the script perfectly, but they still struggle to see the rare, scary storms that matter most. The researchers suggest that until AI can get better at spotting these rare, high-stakes events, we shouldn't let them drive the spaceship alone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →