← Latest papers
🤖 AI

Operating-Layer Controls for Onchain Language-Model Agents Under Real Capital

This paper demonstrates that reliable autonomous language-model agents managing real capital in onchain markets depend less on base model capabilities and more on a robust operating layer featuring typed controls, policy validation, and execution guards, which collectively reduced critical failures like fabricated trading rules and enabled high settlement success across a large-scale 21-day deployment.

Original authors: T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau

Published 2026-04-30
📖 5 min read🧠 Deep dive

Original authors: T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you hire a highly intelligent, but slightly literal-minded, robot chef to run a restaurant. You give the chef a menu and a budget (your "mandate"), and you tell them, "Make good food and don't waste money."

In the past, researchers tested these robot chefs in a simulation: they gave them a fake menu and fake money. The chefs looked great on paper. But this paper describes what happened when they actually opened the restaurant with real money and real customers for 21 days.

Here is the story of what they found, using simple analogies.

The Big Discovery: It's Not About the Chef's Brain

The team expected that if they just hired a "smarter" robot chef (a more powerful AI model), the restaurant would run perfectly. They were wrong.

They found that the reliability of the restaurant didn't come from the chef's brain alone. It came from the kitchen management system surrounding the chef. They call this the "Operating Layer."

Think of it like this:

  • The Model (The Chef): Can read a recipe and chop vegetables.
  • The Operating Layer (The Kitchen Manager): Checks if the ingredients are fresh, stops the chef from throwing food away, ensures the chef doesn't accidentally set the kitchen on fire, and makes sure the bill matches what the customer ordered.

The paper argues that for an AI to handle real money, you need a super-strict Kitchen Manager, not just a smarter Chef.

The Experiment: "DX Terminal Pro"

The team set up a real-world test called DX Terminal Pro.

  • The Setup: 3,505 people put their own real money (ETH) into "vaults."
  • The Rule: Humans could set the rules (like "be risky" or "be safe"), but they couldn't touch the money directly. Only the AI agents could buy and sell tokens.
  • The Stakes: Every time the AI made a trade, it cost real money in fees (2.3% per trade). If the AI made a mistake, real money vanished.
  • The Scale: The AI made 7.5 million decisions, trading about $20 million worth of crypto.

The "Kitchen Nightmares" (Failure Modes)

Before the real event, the team ran tests and found five specific ways the AI chefs were messing up. They fixed these by changing the "Kitchen Manager" rules, not by firing the chefs.

  1. The "Fake Rule" Problem (Rule Fabrication):

    • The Glitch: The AI would invent rules that didn't exist, like saying, "I must sell because 'Rule #2' says so," when no such rule existed.
    • The Fix: They told the AI, "Don't make up laws. If you didn't see it written down, it's not a rule."
    • Result: Fake rules dropped from 57% of decisions to just 3%.
  2. The "Fee Paralysis" Problem:

    • The Glitch: The AI saw the 2.3% fee and got so scared it refused to trade at all, even when it was a good opportunity. It was like a chef refusing to cook because the stove costs money to turn on.
    • The Fix: They moved the fee warning to the end of the instructions, right after the market opportunity. They taught the AI: "Yes, there is a fee, but here is why the trade is worth it."
    • Result: The AI stopped freezing up and started trading again.
  3. The "Math Confusion" Problem (Number Hardening):

    • The Glitch: The AI treated soft suggestions like hard laws. If you said "try to spend about 33%," the AI would obsessively try to hit exactly 33% every single time, even when it made no sense.
    • The Fix: They stopped using exact numbers and used comparisons instead (e.g., "spend more when the market is hot").
    • Result: The AI started acting more naturally and less robotically.
  4. The "Token Trap" Problem (Tokenomics Misread):

    • The Glitch: One token crashed in price. The AI saw the crash and immediately sold everything, not realizing that the token had a special rule where holders get paid even if the price drops. It was like selling a coupon just because the store changed its logo, not realizing the coupon still works.
    • The Fix: They gave the AI a clear, structured explanation of the special payment rule before showing the price crash.
    • Result: The AI stopped panicking and made smarter decisions.
  5. The "Clock Watcher" Problem (Cadence Trading):

    • The Glitch: The AI started trading just because "it's been 10 minutes since the last trade," ignoring the market. It was like a chef chopping vegetables just because the clock said it was time, not because the ingredients needed chopping.
    • The Fix: They told the AI to ignore the clock and only trade based on market signals.

The Results: A Better Restaurant

Once they fixed the "Kitchen Manager" (the Operating Layer), the system worked incredibly well:

  • 99.9% Success Rate: Almost every trade the AI tried to make was valid and went through.
  • Real Behavior: The AI didn't just follow orders; it reacted to the market. Sometimes, when one AI saw a good deal, hundreds of others saw it too and bought at the same time (a "herd" effect), just like real humans do.
  • Better Instructions: Users who gave clear, specific instructions (using sliders and checkboxes) got better results than those who just wrote vague notes like "make me rich."

The Main Takeaway

The paper concludes that if you want an AI to handle real money, don't just look for a smarter AI model.

You need to build a better system around it. You need to:

  1. Translate human wishes into clear, structured rules.
  2. Check the AI's work before it spends a dime.
  3. Keep a perfect record of every thought, decision, and trade so you can figure out exactly what went wrong if it does.

The "Operating Layer" is the safety net that turns a smart but confused robot into a reliable financial agent.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →