Bet on Features: Anytime-Valid and Feature-Aware Auditing of Conditional Quantile Forecasters
This paper introduces a distribution-free, game-theoretic framework for continuously auditing black-box conditional quantile forecasters that accounts for feature-dependent information sets, enabling interpretable, anytime-valid detection of miscalibration without relying on i.i.d. assumptions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a referee watching a magic trick performed by a mysterious "Black Box" oracle. This oracle predicts the future, specifically telling you, "There is a 90% chance the next number will be below this line." In the real world, businesses use these kinds of predictions to decide how much stock to keep on shelves or how much money to set aside for risks.
For decades, referees have checked these predictions using a simple rule: "Did the numbers stay below the line about 90% of the time?" If the answer is yes, the oracle gets a passing grade. But here's the problem: this old rule is like checking a weather forecast by only looking at the average temperature over a whole year. It might look perfect on average, but it could be completely wrong on every single Tuesday in July. If the oracle is consistently wrong on specific days, a business relying on it will lose money, even if the "average" looks fine.
The New Game: Betting on the Details
The authors of this paper, Ivane Antonov and his team, propose a new way to referee these predictions. Instead of just counting the total hits, they turn the audit into a continuous betting game.
Imagine the auditor is a gambler who gets to place a bet every single day.
- The Oracle makes a prediction.
- The real number arrives.
- The auditor bets on whether the Oracle was right or wrong, but with a catch: the auditor can only use information they can see before the bet is placed.
If the Oracle is truly fair (calibrated), no matter how the auditor bets, they will never make a fortune. Their wealth will stay flat, like a coin that lands on heads and tails equally. But if the Oracle is secretly cheating—say, always missing the mark on Saturdays or during sales events—the smart auditor can spot this pattern. By betting specifically on those days, the auditor's wealth will start to grow.
The "Feature-Aware" Twist
The paper's biggest discovery is that what you know changes what you can prove.
Think of the auditor's knowledge like a pair of glasses.
- The "Marginal" Glasses (Blurry): These glasses only show the total score. If you wear these, you might miss the fact that the Oracle is terrible at predicting sales on Saturdays. The auditor stays safe (they won't falsely accuse a fair Oracle), but they are blind to the cheating.
- The "Feature-Aware" Glasses (Clear): These glasses let the auditor see specific details, like "Is it Saturday?" or "Is there a promotion happening?" When the auditor wears these glasses, they can spot that the Oracle is failing specifically on Saturdays.
The authors show that if you only use the blurry glasses, you might think the Oracle is perfect, even though it's failing miserably in the real world. But if you use the clear glasses and bet on the specific features (like "Saturday"), the cheating becomes obvious.
What They Found (and What They Didn't)
The team tested this idea in two ways:
- In a Simulation: They created a fake world where they knew the Oracle was either perfect or secretly biased. They found that their "feature-aware" betting strategy could detect the bias quickly, while the old "blurry" method missed it entirely. In these simulations, the new method correctly identified the cheating in the vast majority of cases when the bias was there, and rarely cried "foul" when the Oracle was actually fair.
- In the Real World: They took a popular, cutting-edge AI model called Chronos-2 and tested it against real store sales data (from a dataset called Rossmann) and synthetic data. The result? The AI model passed the old, blurry tests. But when the authors put on their "feature-aware" glasses and bet on things like promotions and Saturdays, the AI's wealth grew rapidly, crossing the threshold to prove it was miscalibrated. The AI was consistently wrong on those specific days, even though it looked fine on average.
What They Rule Out
The paper explicitly argues against the idea that you can just check the "average" performance to be sure a forecast is good. They show that a forecast can look perfect on average while being dangerously wrong in specific contexts. They also rule out the need for the data to be "independent and identical" (i.i.d.), a common assumption in older statistics. Their method works even when the data is messy, changes over time, or depends on what happened yesterday.
How Sure Are They?
The authors are very confident in their math. They proved that if a forecaster is cheating in a way that is visible through the features the auditor can see, the auditor's wealth will grow, and they will eventually catch the cheat. This isn't just a guess; it's a mathematical guarantee.
However, they are careful to note that their method only works if the auditor has the right "glasses." If the auditor doesn't have access to the specific feature that reveals the cheat (like not knowing it's a Saturday), they might still miss it. The method is powerful, but it's only as good as the information the auditor is allowed to use.
In short, the paper suggests that to truly trust a black-box forecaster, you can't just look at the big picture. You have to bet on the details, and you need to know which details matter. If you do, you can catch the cheaters that the old methods let slip by.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.