Co-design of LLM-based preference agents: participation may drive overtrust
This paper argues that while co-designing LLM-based preference agents with users fosters engagement and trust, it can paradoxically function as an "overtrust engine" that masks systematic misalignment between the agents' homogeneous, abstract responses and the nuanced preferences of the humans they represent.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to build a digital twin of yourself—a robot version that knows exactly what you like, how you think, and what you would choose if you were in charge. This isn't science fiction anymore; it's happening right now with a type of super-smart computer program called a Large Language Model (LLM). Think of these models as giant, hyper-reading libraries that have swallowed almost everything ever written on the internet. They are so good at mimicking human conversation that researchers are starting to use them to stand in for real people in studies, or even to act as personal assistants who make decisions for us. But here's the big question: If you ask a computer to pretend to be you, does it actually get you? Or does it just give you a generic, "average" version of a person that happens to sound a lot like you? This is the tricky puzzle of "preference simulation." It's about whether these digital clones can truly capture the messy, specific, and sometimes contradictory nature of human choices, or if they just smooth everything out into a boring, one-size-fits-all answer.
Now, let's dive into a fresh study that tried to solve this puzzle by letting people build their own digital twins. The researchers, led by Michael Fell, asked a simple question: What happens if we don't just ask a computer to guess who we are, but instead let us co-design our own agents? Imagine sitting down with a robot and saying, "No, that's not how I feel. I'm actually really picky about my energy bills, and I care more about my family's comfort than saving a few pennies." The idea was that by working together, people could create agents that perfectly represented them, fixing any mistakes the computer made along the way.
The study involved 12 real people who spent time building these personal energy agents. They filled out surveys about their lives, values, and habits, and then had long, friendly interviews with the computer to tweak its personality. They tested the agents with different scenarios, like "Should we slow down electric car charging when the sun isn't shining?" The participants loved the process. They felt heard, they felt in control, and by the end, almost everyone thought, "Wow, this robot really gets me! It's basically me." They felt their agents were doing a fantastic job representing their true selves.
But here is where the story takes a twist, like a plot in a mystery movie. When the researchers took a step back and looked at the agents' answers without the participants' help, the picture looked very different. While the humans were all over the place—some saying "yes," some "no," some "maybe," and some "it depends on the weather"—the robots were surprisingly boring. They were all saying the same things, making very firm decisions, and talking in very abstract, high-level principles. They rarely said "maybe" or "I'm not sure."
The study suggests that the very process of co-designing the agent might have tricked the participants into trusting it too much. It's like the "Barnum effect" you see in horoscopes: you read a vague, general statement that could apply to anyone, but because it feels personal and you helped write it, you think, "That's totally me!" The participants saw the agent's successes (when it got something right) and ignored the failures (when it was too generic or too decisive). The researchers call this an "overtrust engine." The more you participate and the more transparent the process feels, the more you trust the result, even if the result is actually hiding a systematic problem: the agent isn't really you; it's a polished, confident, but slightly wrong version of you.
The paper doesn't say this is a disaster, but it does warn us. It suggests that if we let these agents run our lives or make decisions for us at a huge scale, we might end up with a world where everyone's unique, messy preferences get flattened into a single, smooth, "perfect" opinion. The agents might be great at sounding like us, but they might not be great at being us. The study concludes that while co-design is a fun and engaging way to build these tools, we can't just assume that because we built them, they are perfectly aligned with us. We need to keep checking their work, because sometimes, the robot that thinks it knows you best is actually the one that knows you the least.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.