← Latest papers
💻 computer science

Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level

This position paper argues that traditional, static auditing methods fail to capture the emergent and evolving harms of personalized generative AI systems, necessitating a shift toward user-centered, interaction-level approaches that recognize harm as an adaptive and pluralistic process.

Original authors: Hannah Cha

Published 2026-08-18
📖 6 min read🧠 Deep dive

Original authors: Hannah Cha

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a conversation with a machine that remembers everything you have ever said, learns your preferences, and slowly reshapes its own personality to fit you perfectly. This is the promise of personalized generative artificial intelligence, a technology now appearing in chatbots and writing assistants that adapt their tone, style, and advice based on your past interactions. Unlike older computer programs that simply chose from a fixed list of answers, these new systems create entirely new responses on the fly, changing their behavior as the relationship with the user deepens. While this adaptability makes interactions feel more relevant, it also introduces a hidden danger: the very thing that makes the system feel helpful to one person might feel hurtful or alienating to another, and these harms often only appear after weeks or months of talking, not in a single test.

Researchers have long tried to find and fix problems in artificial intelligence by running simulations, where they ask the computer the same questions over and over to see if it produces biased or offensive answers. This method works well for finding static errors, like a map that always gets a specific city wrong. However, a new paper by Hannah Cha from Microsoft Research and Stanford University argues that this approach is failing to catch the real harm happening in personalized systems. The study suggests that because these machines evolve with the user, the definition of what is harmful changes over time and varies wildly even between people who share the same background. The researchers found that trying to fix this by making the personalization even deeper could actually make things worse for marginalized groups, forcing them to do extra work to explain their pain or reveal too much private information just to get the machine to behave safely.

The core of the problem lies in how we currently measure harm. Most existing tests assume that a harmful output is a fixed object, like a broken part in a car, that can be identified before the car ever drives on the road. In the world of personalized AI, the researchers argue, harm is more like a conversation that goes wrong. A response that seems harmless to a researcher in a lab might feel deeply offensive to a user because of their specific history with the system. For example, a machine might start offering food recommendations that seem helpful at first, but after months of interaction, the user realizes the suggestions are based on stereotypes about their culture, turning a helpful gesture into a microaggression. These harms emerge from the flow of the relationship, not from a single mistake, meaning that traditional tests which look at the machine in isolation miss the danger entirely.

Furthermore, the paper points out that current methods often treat groups of people as if they all feel the same way. Audits frequently check if a system treats "women" or "Black people" fairly as a whole block. But the researchers show that within any single group, people have different lives, different traumas, and different boundaries. What one person in a community finds offensive, another might find acceptable, and a third might not notice at all. When a system tries to learn from a group's average behavior, it often flattens these complex differences into a single, simplified version of that group. This can lead to a situation where the machine learns the preferences of the most dominant voices in a community and ignores the rest, effectively erasing the unique experiences of those on the margins.

The researchers also explored a tempting solution: what if the machine just learned exactly what each individual user found harmful? While this sounds ideal, the paper argues it creates an unfair burden. To teach a machine what hurts, a user has to keep pointing out every mistake, which requires constant emotional labor. For people whose experiences already differ from the machine's default settings, this means they must do the heavy lifting to correct the system, while others who fit the default get a smooth experience without effort. Additionally, to teach the machine, users might have to reveal sensitive details about their identity or their past hurts, creating a privacy trap where they must choose between enduring harm or surrendering their personal data. This dynamic risks pushing the most vulnerable users away from the technology entirely.

Instead of trying to perfect the machine's ability to guess what users want, the paper proposes a shift in how we design these systems. The author suggests building infrastructure that allows users to speak up about harm as it happens, not just after the fact. This would involve tools where users can flag a response as problematic and explain why, without the system assuming that one person's complaint is the final truth for everyone. Crucially, this system would also allow users to see how others in their community have interpreted similar situations, helping them decide if a shared explanation resonates with their own experience. The goal is to move away from a model where the machine silently learns and the user passively receives, toward a partnership where the user retains control over how the machine adapts.

The researchers acknowledge that this approach is not a simple fix. It requires building trust in systems that are often opaque, and it raises difficult questions about how to handle conflicting opinions within a community. If one person says a comment is harmful and another says it is fine, the system must be able to hold both truths without forcing a single answer. The paper does not claim to have solved these technical challenges, but it insists that the current path of static testing and deep personalization is insufficient. By recognizing that harm is a fluid, evolving process shaped by the user's history and context, we can begin to design AI that respects the complexity of human experience rather than trying to simplify it away. The ultimate finding is that protecting users in a personalized world requires giving them a voice in the conversation, rather than just hoping the machine learns the right lesson on its own.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →