DAR: Deontic Reasoning with Agentic Harnesses
This paper introduces Deontic Agentic Reasoning (DAR), an agentic framework that enables large language models to interact with statutes on demand to improve performance on complex deontic reasoning tasks, though it notes that these benefits are not uniform and can lead to degraded numerical performance and higher token consumption in weaker models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but the only clue you have is a massive, 1,000-page rulebook written in dense legal language. Your job is to find the specific rule that applies to a specific person's situation and calculate the result.
This paper is about testing how well different "detectives" (AI models) can solve these mysteries when they are allowed to use a special toolset versus when they are just handed the whole book at once.
The Two Ways to Solve the Mystery
The researchers compared two different approaches:
- The "Big Stack" Approach (Direct Reasoning): Imagine handing the detective the entire 1,000-page book, the case facts, and the question all at once. The detective has to read through the whole thing in their head, find the right page, and write the answer immediately.
- The "Search & Dig" Approach (Deontic Agentic Reasoning or DAR): Instead of giving the detective the whole book, you put the book on a shelf in a room and give them a set of tools (like a magnifying glass, a search engine, and a calculator). The detective has to walk up to the shelf, use the tools to find the specific page they need, read it, and then come back to write the answer. They can do this as many times as they need.
The Big Discovery: It Depends on Who You Ask
The researchers tested this with two types of detectives: Super-Geniuses (top-tier, expensive AI models) and Apprentices (smaller, open-source AI models).
1. The Super-Geniuses Thrived
When the Super-Geniuses were allowed to use the "Search & Dig" tools, they got much better at solving the puzzles.
- Why? They are smart enough to know what to look for. They could say, "I need to find the section on tax deductions," use the
greptool to jump straight there, and ignore the rest of the book. - The Result: Their accuracy jumped significantly (sometimes by 15–30%). They used the tools efficiently to find the right rules and calculate the answer correctly.
2. The Apprentices Struggled (and Got Worse)
When the Apprentices were given the same tools, they actually performed worse than when they were just handed the whole book.
- Why? They got confused by the freedom. Instead of stopping once they found a clue, they kept digging, reading irrelevant pages, and getting stuck in loops. They spent a huge amount of time (and computer "energy") confidently writing long, wrong answers.
- The Result: Their accuracy dropped dramatically (sometimes falling to near zero). They consumed up to 4 times more computer resources just to get a worse answer.
The "Confidence Amplifier" Metaphor
The paper suggests that for the weaker models, the tools acted like a confidence amplifier.
- Imagine a student who doesn't know the answer to a math problem. If they just have to write it down, they might guess and move on.
- But if you give them a calculator and a textbook and say, "Go find the answer," they might start frantically flipping pages, doing random calculations, and convincing themselves they are right, even though they are wrong. The tools made them think they were working hard, but they were just spinning their wheels.
The Cost of Being Wrong
The paper highlights a major cost issue.
- The Super-Geniuses were efficient. They used the tools to get the right answer quickly.
- The Apprentices were wasteful. They used the tools to generate massive amounts of text, burning through computer resources (tokens) without improving their accuracy. In some cases, they ran out of time before they could even finish their thought.
The Bottom Line
The paper concludes that giving an AI a "toolbox" to look up rules isn't a magic fix for everyone.
- If the AI is already very smart, the toolbox helps it shine.
- If the AI is still learning, the toolbox can actually make it overconfident, wasteful, and less accurate.
The researchers warn that we shouldn't assume giving an AI more tools automatically makes it better at complex tasks like tax law or immigration rules. For weaker models, it might just make them more expensive and more prone to confident mistakes.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.