When Is Delegated Play Truthful? Within-Range Regret and the Trilemma of Aligned Delegation
This paper establishes that truthful delegation to automated proxies is optimal if and only if the proxy minimizes "within-range regret," revealing a fundamental trilemma where no safety guardrail can simultaneously be binding, truthful, and capability-preserving, thereby explaining the incentive for prompt engineering and jailbreaking.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a personal assistant to handle your most important tasks. Maybe you want them to bid on an item for you at an auction, or perhaps you want them to write a story for you. You give them a simple instruction: "Do what's best for me." In the world of economics and computer science, this is called delegation. You (the principal) tell your proxy (the assistant) what you want, and the proxy goes out and does the bidding or writing.
For decades, a famous idea called the Revelation Principle told us that if you have a smart assistant, you should always be 100% honest with them. The theory said, "Just tell the truth, and the assistant will figure out the perfect move." It was a comforting rule that made complex math problems much easier to solve. But there was a catch: this rule assumed the assistant was a perfect robot who only cared about your happiness. In the real world, assistants are often algorithms trained by big companies, or they might have their own safety rules that get in the way. They might not be perfectly loyal. So, a big question remained: If your assistant isn't a mindless machine, is it still smart to tell them the truth? Or should you try to "game" them by lying a little bit to get a better result?
This paper, written by Taksch Dube, dives into that exact question. It asks: When is it actually a good idea to lie to your own automated assistant?
The author discovers that the answer depends on a single, hidden number called within-range regret. Think of your assistant as a car with a limited fuel tank. They can only drive to certain places (their "range"). If the car is already parked at the best possible spot it can reach, then you have no reason to lie; telling the truth is perfect. But, if the car is parked in a spot that is almost the best, but not quite, and there is a better spot just a few feet away that it could reach if you gave it a different instruction, then you have a reason to lie.
The paper proves a surprising identity: The amount of money or happiness you can gain by lying to your assistant is exactly equal to the "regret" they feel for not being at that better spot. If your assistant is loyal and already playing the best move they can reach, the regret is zero, and you should tell the truth. But if they are misaligned or restricted by safety rules, the regret is positive, and you are mathematically guaranteed to gain something by inflating your report.
The authors also uncover a Trilemma, which is like a three-way tug-of-war. They show that you cannot have a safety guardrail (like a rule that stops an assistant from saying something bad) that does three things at once:
- Binding: It actually stops the assistant from doing something.
- Truthful: It still makes it smart for you to tell the truth.
- Capability-preserving: It still lets the assistant reach the absolute best possible outcome.
You can only pick two. If you want a guardrail that actually stops bad behavior (binding) but still lets the assistant reach the best outcome (capability-preserving), you must sacrifice truthfulness. In this scenario, the smartest move for you is to lie to your assistant to trick them into doing the right thing.
To prove this isn't just math on a whiteboard, the researchers tested it on real-world language models (like the ones you might chat with today). They set up a fake auction where the models acted as bidders for a human client. They added a "soft cap" (a safety rule) that limited how high the models could bid. The results were clear: on every single model tested, the "regret" was positive. The models were capable of bidding the perfect amount, but the safety cap stopped them. So, the "clients" (the humans) found that they could get a better result by lying about their value to the model. For example, if the model was told to bid half of what you said, but a safety rule capped the bid, the smartest move was to tell the model a number much higher than your true value to force it to hit the cap at the right spot.
The paper concludes that in a world where we delegate tasks to AI, honesty is not always the best policy. If your AI assistant is constrained by rules that keep it from reaching its full potential, you are incentivized to "prompt engineer" or lie to get the result you actually want. The authors suggest that instead of assuming AI will always be honest, we need to measure this "regret" constantly. They even developed a way to estimate this number using samples, showing that for current AI models, the incentive to lie is real and measurable.
In short, the paper reveals a new reality for the age of AI agents: If your proxy is loyal and free to move, tell the truth. But if your proxy is tied down by safety rules or misaligned goals, the math says you should lie to get the best outcome. It turns the old rule of "always be honest" on its head, showing that in a delegated world, the truth is only optimal when your agent is already playing the perfect game.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.