Beyond Localization: Recoverable Headroom and Residual Frontier in Repository-Level RAG-APR
This paper investigates the recoverable performance gains and remaining limitations of repository-level RAG-APR systems beyond localization by evaluating Oracle Localization, candidate diversity, context injection, and interface design on SWE-bench Lite, revealing that while stronger localization and specific post-localization levers improve repair rates, a significant residual frontier persists due to inherent system and prompt-level constraints.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to fix a massive, ancient library (a software repository) where a specific book has a typo that breaks the whole story. You have a team of three very smart, but slightly different, librarians (the AI repair systems: Agentless, KGCompass, and ExpeRepair).
For a long time, everyone thought the only way to fix the library was to get the librarians better at finding the exact page with the typo. This is called "Localization."
This paper asks a new, crucial question: "Okay, let's pretend we magically give the librarians the exact page number where the typo is. Now that they know where to look, what stops them from actually fixing it?"
The authors ran a series of experiments to find out what happens after the location is found. Here is what they discovered, using some simple analogies:
1. The "Magic Map" Experiment (Oracle Localization)
The researchers gave the librarians a "Magic Map" that pointed directly to the typo.
- The Result: The librarians got much better at fixing the books, but they still failed more than half the time.
- The Analogy: Imagine you are told exactly which brick in a wall is broken. You can now see the problem clearly, but you might still fail to fix it because you don't have the right tools, you don't know the right mortar recipe, or you just can't reach the spot.
- Key Finding: Knowing where the problem is helps, but it doesn't solve the whole puzzle. The real struggle happens after you find the spot.
2. The "Lottery Ticket" Experiment (Search Headroom)
The researchers asked: "If we ask the librarians to try 10 different ways to fix the typo (instead of just one), will they get better?"
- The Result: Yes, but only a little bit. The first few attempts were the most helpful. By the 5th attempt, they had already found almost all the "good" ideas they could. Trying 10 times instead of 5 didn't help much more.
- The Analogy: It's like buying lottery tickets. The first few tickets you buy give you a decent chance of winning. Buying 10 tickets instead of 5 doesn't double your luck; you just run out of "winning" tickets pretty quickly. The librarians ran out of good ideas very fast.
3. The "Extra Clues" Experiment (Added Context)
The researchers tried giving the librarians extra notes from the other librarians. For example, they gave Agentless a note from KGCompass saying, "Hey, I saw a similar problem in a different book."
- The Result: This helped! The librarians fixed more books when they had these extra, relevant clues.
- The Catch: It wasn't just about having more paper to read. If they gave them extra paper with useless junk on it, the librarians got confused and fixed fewer books.
- The Analogy: It's like cooking. If you give a chef the exact recipe (the right clue), they make a great meal. If you give them the recipe plus a 50-page manual on how to grow wheat (too much context), they get overwhelmed. But if you give them a tip from a master chef on how to chop onions (relevant extra context), the meal gets even better.
4. The "Unsolvable Frontier" (The Residual Problem)
Finally, the researchers asked: "If we combine all the librarians' best attempts, the Magic Maps, and the extra clues, how many books are still broken?"
- The Result: A surprisingly large number. Even after doing everything right, there was a "Frontier" of problems that no one could fix.
- The Analogy: Imagine a group of detectives solving crimes. Even if they have the perfect crime scene photo, the best forensic tools, and all the witness statements, there are still some cases where the culprit is just too clever, or the evidence is too messy.
- The Reality: The remaining broken books usually had very specific, tricky errors (like a missing ingredient in a recipe or a confusing instruction). The current AI tools just aren't smart enough to handle those specific, messy details yet.
The Big Takeaway
For a long time, the AI repair community thought: "If we just get better at finding the bug, we will fix everything."
This paper says: "Nope. Finding the bug is only step one. Once we find it, the real hard work begins."
To fix software in the future, we need to stop just focusing on "finding" and start focusing on:
- How to use the evidence once we find it.
- How to design better tools for the AI to use.
- Accepting that some problems are currently too hard for our AI to solve, no matter how much we tweak the search.
In short: We found the needle in the haystack. Now we need to figure out how to actually sew the hole in the haystack.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.