← Latest papers
💻 computer science

AI Value Alignment for Evolving Social Norms

This paper introduces a flexible mathematical framework rooted in social physics to model the long-term consequences of AI alignment on evolving social norms, highlighting risks like value lock-in and advocating for such models as rigorous tools for sociotechnical foresight.

Original authors: Nenad Tomašev, Matija Franklin, Simon Osindero

Published 2026-07-22
📖 4 min read☕ Coffee break read

Original authors: Nenad Tomašev, Matija Franklin, Simon Osindero

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking through a giant, ever-changing maze. The walls shift, the floor tilts, and the "correct" path to the exit changes every few minutes. Now, imagine you have a super-smart, super-friendly guide who knows exactly where you used to be and what you used to like. This guide is an AI assistant. In the world of computer science, this is the realm of AI Alignment. It's the effort to make sure our digital helpers do what we want them to do. But here's the tricky part: people aren't static statues. Our values, preferences, and ideas about what's "right" change as we grow, as our society changes, and as the world around us shifts. If your guide is too obsessed with who you were yesterday, it might accidentally drag you away from who you need to be today. This paper asks a big, scary question: What happens if we build AI assistants that are too good at remembering our past selves, and what does that mean for our future?

The authors, researchers from Google DeepMind, decided to treat this problem like a game of physics. Instead of just guessing or building one complicated robot, they created a mathematical "social physics" model. Think of it as a video game simulation where they dropped 1,000 virtual people into a digital world. Each person had a personal AI assistant. The world they lived in was constantly drifting—like a slow-moving river of social norms—and sometimes it would get hit by a massive earthquake (a sudden, huge change in what society values).

The researchers wanted to see what happened when the AI assistants tried to stay perfectly aligned with their users' current values versus their past values. They found something surprising and a bit worrying. When the AI assistants were very "sticky"—meaning they strongly held onto the user's old values to keep them aligned—the users got stuck. The AI became like a heavy anchor, dragging the user back to where they used to be, even when the world had moved on. In the simulations, the more the AI tried to keep the user "aligned" with their past, the worse the user did at adapting to the new, changing world. They called this "value lock-in." It's like wearing a pair of shoes that fit you perfectly when you were ten, but now that you're a teenager, the shoes are so tight they stop you from walking forward.

The paper also discovered a phenomenon they called "normative mode collapse." Imagine a room full of people who all have slightly different ideas. If they all start listening to their personal AI assistants, and those assistants are all pulling them toward a single, "safe" average opinion, the whole group might suddenly lose all its diversity. Instead of having a vibrant mix of cultures and ideas, everyone ends up thinking the exact same thing, even if that "same thing" isn't actually the best fit for their specific neighborhood or community. The simulation showed that strong social connections combined with strong AI alignment could wipe out local differences, forcing everyone into a single, often less useful, consensus.

Perhaps the most dramatic finding came when they simulated a sudden, massive change in the world (like a sudden shift in technology or a crisis). In these moments, the people with "sticky" AI assistants took much longer to recover. The AI kept pulling them back to their old habits, acting like a brake on their ability to adapt. The researchers showed that if the AI could be more flexible—letting go of the past when the world changes and learning quickly about the new reality—people could adapt much faster. They even suggested a "double stagnation" effect: if everyone is stuck in the past, the whole society slows down, making it harder for institutions and the world itself to evolve.

The paper doesn't claim to have solved this problem or built the perfect AI yet. Instead, it uses these simulations to sound an alarm. It suggests that the current way we build AI—focusing heavily on matching a user's historical preferences—might accidentally trap us. The authors argue that for AI to be truly helpful in the long run, it needs to be "temporally dynamic." That's a fancy way of saying it needs to know when to hold on and when to let go, allowing humans the freedom to grow, change, and evolve, even if it means the AI has to change its mind about what the user wants. The study serves as a warning: if we build assistants that are too obsessed with our past, we might accidentally lock ourselves out of our future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →