Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
This paper introduces Harness-IF, a novel benchmark that evaluates coding agents' true instruction-following capabilities across diverse surfaces by distinguishing genuine compliance from coincidence using a new metric called Against-Prior Accuracy (AP-Acc), revealing that existing benchmarks overstate performance and that rule precedence does not strictly follow prompt depth.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot chef to cook a complex meal. You don't just shout one command like "Make pasta!" at the very end. Instead, you have a whole stack of instructions: a rulebook on the wall (the system prompt), a note on the recipe card (the project file), a description of how the knife works (the tool description), and finally, your specific request to the chef (the user instruction). For a long time, scientists testing these AI chefs only cared about one thing: did the chef serve the pasta? If the plate looked good, they gave a gold star. But they never checked which instructions the chef actually listened to. Maybe the chef ignored the rulebook about "no garlic" because they just happened to love pasta without garlic anyway. This is the tricky corner of artificial intelligence we are exploring: Instruction Following. It's not just about getting the right answer; it's about proving the robot followed the specific rules you gave it, even when those rules go against what the robot would have done naturally.
This paper, titled "Harness-IF," is like a detective agency for AI chefs. The researchers realized that existing tests were too easy because they only looked at the final dish. They wanted to know if the robot was actually obeying the rules or just getting lucky. So, they built a new, super-detailed test called Harness-IF. Instead of just asking, "Did you make the pasta?", they set up 60 realistic coding scenarios where the AI had to follow 256 different tiny rules. These rules were hidden in different places in the robot's "brain"—some in the system prompt, some in project files, some in tool descriptions, and some in the user's direct chat.
Here is the big surprise the detectives found: AI models are much better at following rules that match what they would have done anyway than rules that go against their natural habits. The researchers call this the "Against-Prior" test. They found that when a rule forced the AI to do something different from its default behavior, the AI's success rate dropped significantly. On average, the models were about 5.8 percentage points worse at following rules that went against their grain. It's like a student who gets an A on a math test because they love math, but gets a C when the teacher asks them to solve a problem using a method they hate. The paper suggests that if we only look at the final grade (the task success), we are tricking ourselves into thinking the student is a math genius, when they might just be coasting on their natural talent.
The study also looked at where the rules were placed. You might think the rule spoken last (the user instruction) would be the most important, like the last person to speak in a meeting usually wins. But the data showed something different. Rules found in the system prompt, project files, and user instructions all tied for the top spot in importance, while rules buried in tool descriptions or skill descriptions were often ignored. It turns out the AI pays attention to the "big picture" rules more than the nitty-gritty details of how a specific tool works.
In short, this paper doesn't claim to have solved the problem of making perfect AI. Instead, it gives us a better ruler to measure how well AI actually listens. It shows us that current AI models are great at doing what they want to do, but they still struggle when we ask them to do something they wouldn't naturally choose. By separating "lucky compliance" from "true obedience," the researchers hope we can build better tests to make our AI helpers more reliable, not just more successful-looking.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.