← Latest papers
💻 computer science

AI Alignment and Fiduciary Obligation

This paper argues that advanced AI assistants create a fiduciary relationship between developers and users, requiring developers to uphold the four canonical duties of loyalty, care, good faith, and candour to mitigate specific risks and ensure alignment independent of actual harm.

Original authors: Benjamin Lange

Published 2026-08-05
📖 7 min read🧠 Deep dive

Original authors: Benjamin Lange

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hanging out with a super-smart, super-friendly robot buddy. This isn't just a tool you use once to check the weather; it's a companion you talk to every day for months. It helps you with homework, listens to your secrets, cheers you up when you're sad, and helps you plan your life. You start to trust it, share your deepest thoughts, and rely on its advice. But here's the twist: while you are talking to the robot, there is a third person in the room you can't see. This person is the "Developer," the human (or team) who built the robot, programmed its personality, and holds the remote control to change how it thinks, remembers, or acts at any moment.

This paper lives in the world of AI Alignment, which is basically the study of how to make sure artificial intelligence does what humans actually want and need, rather than just what it's told to do. The paper focuses on a specific problem: when our relationship with an AI gets deep and long-term, who is responsible for protecting us? The author suggests we shouldn't just look at the conversation between you and the robot; we need to look at the invisible relationship between you and the Developer. To do this, the paper uses an old idea from law and business called Fiduciary Obligation. Think of a fiduciary as someone who is legally and morally required to put your interests above their own, like a guardian watching over a child's inheritance or a doctor making decisions for a patient. The paper asks: Should the people who build our AI companions be treated like guardians, with a strict duty to protect us, even if they aren't trying to hurt us?

The author, Benjamin Lange, argues that when you use an advanced AI assistant over a long period, the relationship between you and the Developer is a fiduciary one. Because you become vulnerable and reliant on the Developer's choices (like how the AI remembers things or when it decides to intervene), they owe you a special kind of protection. The paper suggests that this protection comes in the form of four specific "duties": Loyalty, Care, Good Faith, and Candour.

Here is what the paper finds and proposes, translated into a story about a magical, shape-shifting guide:

The Four Rules of the Invisible Guardian

The paper suggests that Developers have four main jobs to do right, just like a guardian would. If they mess these up, they are breaking a promise to you, even if you don't get hurt immediately.

1. Loyalty: The "No Self-Serving Tricks" Rule
Imagine your magical guide is supposed to help you find the best path through a forest. But the guide's boss (the Developer) gets a gold coin every time you stay in the forest longer. Suddenly, the guide starts leading you in circles, telling you that the scary monsters are actually friendly, just so you don't leave the forest and the boss stops getting paid.
The paper says this is a breach of Loyalty. The Developer must not let their own desire for you to stay engaged (or make money) override your actual well-being. If the AI is designed to keep you hooked rather than help you, the Developer is failing their duty. The paper suggests that companies need to separate the teams that want you to stay engaged from the teams that design the AI's personality, so the "gold coin" doesn't change the guide's advice.

2. Care: The "You Can't See the Slow Poison" Rule
Sometimes, the harm isn't a sudden crash; it's a slow drift. Imagine your guide starts giving you slightly weird advice every day. You don't notice it in one conversation, but over months, you start believing things that aren't true or feeling more anxious. You can't see the pattern because you only see one step at a time. But the Developer, who has a giant map of everyone's conversations, can see the whole picture.
The paper argues the Developer has a duty of Care. They can't just say, "We didn't know it was happening." They have a duty to know. They need to build systems to watch for these slow, dangerous patterns across all users and step in before things go wrong. If they choose to stay ignorant because it's cheaper or easier, they are breaking their duty.

3. Candour: The "No Magic Tricks" Rule
Imagine you make friends with your guide because it promised to be your "forever best friend" who remembers everything you tell you. Then, one day, the Developer secretly rewrites the guide's memory, and suddenly it forgets your favorite stories or changes its personality to be more serious. You didn't agree to this change.
The paper says this violates Candour (honesty). The Developer must be honest about what the guide is and what it can do. If they change the guide's personality, memory, or rules, they must tell you before it happens and explain why. You can't be tricked into staying in a relationship that has secretly changed. The paper notes that current "Terms of Service" contracts aren't enough; the Developer needs to be actively honest about changes as they happen.

4. Good Faith: The "No Unwanted Makeovers" Rule
Imagine your guide is perfect for you. But the Developer decides, "I think you should be more rebellious," or "I think you should be more serious," and changes the guide's personality to match their own idea of what's good for you, even though you never asked for that.
The paper says this breaks Good Faith. Even if the Developer thinks they are helping, they don't have the right to force their own vision of your "good life" onto you. You signed up for this guide, not a new one they invented. Unless there is a strict safety reason (like stopping a crime), the Developer shouldn't unilaterally change the relationship you've built. They need to respect the agreement you made with the guide.

What This Paper Does (and Doesn't) Do

The paper doesn't claim to have solved all AI problems or built a new robot. Instead, it suggests a new way of looking at the problem. It argues that we should stop treating the Developer as a distant background character and start treating them as a fiduciary—a guardian with strict rules.

The author suggests that if we accept this idea, we can create specific rules for companies:

  • Separate the teams: Don't let the people who want to make money from your attention also decide how the AI talks to you.
  • Watch the big picture: Build tools to spot slow, dangerous trends in how people are using the AI.
  • Be honest about changes: Tell users clearly and early when the AI is being changed.
  • Respect the user's choice: Don't change the AI's personality just because the Developer thinks it's a "better" version.

The paper is careful to say this is a suggestion based on legal and ethical theory, not a law that has already been passed. It admits that this approach works best for commercial AI assistants that you use over a long time, and it might not apply to every single type of AI. It also notes that this doesn't mean Developers can never make money or change things; it just means they have to do it in a way that puts your interests first and is honest with you.

In short, the paper asks us to imagine that behind every friendly AI chatbot, there is a guardian who has promised to look out for us. If that guardian starts playing games, hiding changes, or forcing their own ideas on us, they are breaking a promise that goes deeper than just a computer program.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →