When benchmark inferences do not compose: Projectibility in AI evaluation
This paper argues that valid AI benchmark inferences do not automatically compose into warranted consequential claims due to misalignments in endpoints, assumptions, and dependencies, proposing a "projectibility audit" framework to diagnose these unsupported logical joins.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to figure out if a new robot chef is ready to run a busy restaurant. You can't just watch it chop onions in a test kitchen and assume it will handle a full dinner service. You have to ask: Will it cook the same way when the ingredients are different? Will it still be good if a human manager checks the food? Will it survive a week of service, or just one hour? This is the heart of AI evaluation, a field where scientists try to measure how smart or capable artificial intelligence really is.
To do this, researchers often use benchmarks. Think of a benchmark as a standardized test, like a driver's ed exam. If a car passes the test, it gets a license. But here's the tricky part: passing a test doesn't automatically mean the car can drive safely in a snowstorm, on a bumpy dirt road, or with a nervous passenger. In the world of AI, we often see a chain of reasoning: "The AI passed the test, so it has 'general reasoning' skills, so it can do this specific job, so it's safe to use." The big question is: Does the evidence from step one actually connect to step two, and does step two connect to step three? This paper explores why that connection often breaks, even when every single step looks perfect on its own.
The Broken Chain of Logic
Imagine you are building a tower out of blocks. You have a block that is perfectly square and sturdy (a benchmark result). You have another block that is also perfectly square and sturdy (a local study showing the AI works in a specific office). You might think, "If I stack them, I have a tall, stable tower!" But what if the top of the first block is slightly rounded, and the bottom of the second block is slightly flat? They look like they fit, but they don't actually touch. The tower wobbles and falls.
This is the main problem Brett Reynolds is pointing out in this paper. He calls it the non-composition principle. In simple terms, just because you have two pieces of evidence that are both "warranted" (supported by good data) doesn't mean you can glue them together to make a bigger, stronger claim. The paper argues that when we try to chain these pieces of evidence together—like going from "AI passed a test" to "AI is safe for real-world use"—we often miss the tiny gaps where the logic falls apart.
The "Legal Assistant" Puzzle
To show how this happens, the author creates a story about a law firm trying to use a new AI tool.
- The Test: A developer shows that their AI model gets 90% of legal questions right on a specific test.
- The Human Check: The law firm shows that when their human lawyers review drafts, they catch 99% of the mistakes.
- The Big Leap: The firm assumes that if they combine these two, the final legal memos will be 99.9% perfect.
The paper says: Stop right there. This logic is broken.
Why? Because the "Test" and the "Human Check" are talking about different things. The test might only check if the AI can find a quote from a book. But the human lawyers might be checking if the AI missed a crucial law that wasn't in the book. If the AI misses that law, the human lawyer might not catch it because they are only looking for the wrong kind of mistake. The two studies are like two people looking at a puzzle from different angles; they both see a perfect picture, but they aren't looking at the same picture.
The paper simulates this with numbers. In one scenario, the AI does great on "Termination Clause" requests (a specific type of legal question) but fails miserably on "Reasonable Notice" requests. If you just average the scores, you might think the AI is great overall. But if the law firm needs to use it for both types of questions, the "average" score is lying to them. The AI is actually a disaster for half the job.
The "Magic Mean" Trap
The paper also warns about a statistical trick called aggregation. Imagine you have a class of students. Half of them get 100% on a math test, and the other half get 0%. The average score is 50%. That sounds okay, right? But if you need every student to pass to run a school bus, that average is useless.
In AI, a "stable mean" (a steady average score) can hide a lot of trouble. The author shows that an AI's score might stay the same on average, but the types of mistakes it makes could be getting much worse. Maybe it used to make small typos, but now it makes dangerous legal errors. If you only look at the average, you miss the fact that the "worst-case" scenarios are getting much worse. The paper uses a simulation to show that you can have a perfectly stable average score while the AI is actually failing in the most critical ways.
What This Means for the Future
So, what does the paper suggest we do? It proposes a "Projectibility Audit."
Think of this as a checklist for anyone trying to use AI. Before you say, "This AI is ready for the real world," you have to answer specific questions:
- Do the blocks fit? Does the "Test" actually cover the same types of questions as the "Real World" job?
- Are the conditions the same? Did the test happen in a quiet room, while the real job is in a noisy, chaotic office?
- Did we lose information? Did we throw away the details about which questions were hard, so we only see the average?
The paper doesn't say AI is bad or that we can't use it. It says we need to be much more careful about how we connect the dots. We can't just assume that because an AI passed a test, it will pass a job. We have to check the connection between the test and the job, and the job and the final result.
The Bottom Line
This paper is a reality check for the AI world. It suggests that we are often too quick to stack our evidence blocks, assuming they fit perfectly when they actually have gaps. The author uses simulations and logical arguments to show that support for one step does not automatically transfer to the next step.
If a developer says, "Our AI is 90% accurate," and a company says, "Our lawyers catch 99% of errors," the paper tells us: Don't multiply those numbers yet. You have to check if the AI is making the same kind of errors the lawyers are looking for. If you don't, you might build a tower that looks tall but collapses the moment you put a real load on it. The paper doesn't claim to have solved the problem of AI safety, but it gives us a better map to find where the cracks in our logic are hiding.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.