← Latest papers
💻 computer science

Contracting for LLM Delegation: Moral Hazard in Technology and Effort Choice

This paper extends the Principal-Agent framework to model LLM delegation, demonstrating that optimal linear contracts with threshold-based technology switching effectively align incentives between principals and agents who jointly select models and effort levels, a finding validated through theoretical derivation and empirical calibration on MATH and MMLUPro benchmarks.

Original authors: Nanda Kishore Sreenivas, Kate Larson

Published 2026-08-20
📖 6 min read🧠 Deep dive

Original authors: Nanda Kishore Sreenivas, Kate Larson

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the rapidly evolving world of artificial intelligence, a quiet revolution is taking place where computers are no longer just answering questions but are being asked to do the work. Instead of a human typing a prompt and waiting for a single reply, we are seeing the rise of autonomous agents—software programs that can break down complex tasks, make decisions, and execute steps on their own. Imagine a law firm asking an AI to review a contract, or a software team asking it to fix a bug. The human provides the goal, but the machine decides how to get there. This shift creates a new kind of relationship, one that economists have studied for decades: the principal-agent dynamic. In this arrangement, one party, the principal, hires another, the agent, to perform a task. The problem arises when the principal cannot see exactly what the agent is doing. If the agent is paid based on how much work it appears to do, it might waste time or choose the cheapest, least effective tools just to save money, leaving the principal with a poor result. This hidden behavior is known as moral hazard, and it becomes especially tricky when the "worker" is a sophisticated AI that can choose between different brainpower levels and decide how hard to think about a problem.

Researchers at the University of Waterloo have stepped into this emerging landscape to figure out how to pay these autonomous agents fairly and effectively. They focused on a specific scenario where an AI agent must choose two things: which large language model to use as its brain, and how much computational effort to spend on a task. Think of the models as different types of engines; some are small, cheap, and fast but might struggle with hard problems, while others are massive, expensive, and powerful but cost a fortune to run. The agent also decides on the "effort," which in the world of AI means how many tokens, or units of text, it generates to solve the problem. The researchers found that the quality of the answer improves as the agent spends more effort, but only up to a point. Like filling a bucket, the first few gallons of water make a huge difference, but after a while, adding more water just overflows the bucket without making it any fuller. This pattern, known as a saturating function, means that spending more money on effort eventually yields diminishing returns.

The core of the study was to design a payment plan, or contract, that would encourage the agent to make the right choices without the principal needing to watch every move. The researchers proposed a simple linear contract: the agent gets a percentage of the value of the final result. If the result is good, the agent gets a bigger slice of the pie. The team mathematically proved that under this system, the agent would naturally switch from a cheaper, weaker model to a more expensive, powerful one only when the reward share crossed a specific threshold. Below that line, the agent sticks to the cheap model to save costs. Above it, the potential for a better result makes the expensive model worth the price. This switching point is not random; it is a precise mathematical boundary determined by the costs and capabilities of the available models. The researchers also showed that the principal can calculate the perfect percentage to offer to maximize their own profit, balancing the cost of the reward against the quality of the work produced.

To see if this theory held up in the real world, the team tested it using actual AI models on two difficult challenges: advanced mathematics and complex general knowledge questions. They gathered data on how different models performed as they were allowed to use more and more computational effort, fitting their real-world performance to the theoretical curves. They then set up a simulation where an AI agent and a human principal learned to interact over thousands of rounds. The agent used a learning algorithm to figure out which model and effort level worked best for different payment offers, while the principal learned which payment percentage yielded the best results. The results were striking. Both the agent and the principal, starting with no prior knowledge of the models' hidden capabilities, eventually converged on strategies that matched the researchers' theoretical predictions almost perfectly. The agents learned to switch models at the exact point the math suggested, and the principals learned to offer the exact reward share that maximized their returns.

The study also explored what happens when models require a "warm-up" period, a phase where they must spend a certain amount of effort just to get started before they can produce any useful answer. This is common in complex reasoning tasks where the AI needs to think for a while before it can write a solution. The researchers found that this initial cost changes the math slightly, pushing the point at which it becomes worth switching to a more powerful model even higher. Even with this added complexity, the learning algorithms in the simulation still found the optimal path, confirming that simple, transparent contracts can guide complex, autonomous systems toward efficient behavior. The work suggests that as we move toward a future where AI agents manage other AI agents, we do not need complicated, opaque systems to keep them honest. A straightforward agreement based on the quality of the outcome is enough to align the interests of the human and the machine, ensuring that the right tools are used and the right amount of effort is applied.

This research does not claim to have solved every problem in the economics of artificial intelligence. The authors note that their model assumes the principal can easily verify the final result, which might not always be true in the real world. They also point out that future work could explore situations where the value of the task is hidden from the agent, creating a deeper layer of uncertainty. However, the findings provide a solid foundation for understanding how to structure deals in an age of autonomous delegation. By treating the choice of technology and the allocation of effort as a single, coordinated decision, the study offers a clear path forward for building markets where AI agents can work efficiently without constant human supervision. The message is one of cautious optimism: with the right incentives, the complex machinery of artificial intelligence can be guided to work for us, not just for itself.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →