← Latest papers
💬 NLP

Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale

This paper introduces a novel framework for black-box auditing of personalization algorithms using generative AI agents with fixed personas to enable scalable, counterfactual analysis of how user attributes influence content delivery, demonstrating through a large-scale study on X that the platform's algorithm amplifies polarizing content and responds differently to demographic signals depending on user ideology.

Original authors: Alessandro Morosini, Sarah H. Cen, Andrew Ilyas, Hedi Driss, Aleksander Mądry, Chara Podimata

Published 2026-07-01
📖 5 min read🧠 Deep dive

Original authors: Alessandro Morosini, Sarah H. Cen, Andrew Ilyas, Hedi Driss, Aleksander Mądry, Chara Podimata

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking into a giant, invisible library where the books on the shelves change depending on who you are. If you look like a young person, you see different books than if you look like an older person. If you seem to live in a city, you see different books than if you seem to live in the country. This is how social media algorithms work: they curate what you see based on your profile and your past behavior.

The problem is that these libraries are "black boxes." We can't see the librarian's rulebook. We can't ask, "Why did you show me this?" We can only watch what happens.

This paper introduces a clever new way to peek behind the curtain using AI agents (computer programs powered by advanced AI) instead of real humans.

The Problem with Old Methods

Previously, researchers tried to audit these libraries in two ways, both of which had flaws:

  1. Real Humans: They hired real people to use the app. This was realistic, but expensive, hard to control, and impossible to scale up. You can't easily ask, "What would happen if this 25-year-old was actually 55?" without finding a new human.
  2. Scripted Bots (Sock Puppets): They created fake accounts controlled by simple computer scripts. These were cheap and scalable, but they were "flat." They acted like robots following a strict list of rules (e.g., "If you see a cat, click like"). They didn't think or react naturally, so the library didn't treat them like real people.

The New Solution: "AI Actors"

The authors created a framework using Generative AI agents. Think of these not as robots, but as actors in a play.

  • The Script (Persona): Before the play starts, each actor is given a detailed character script. This script defines who they are: their age, gender, location, education, and political beliefs. This script is based on real data from the US Census and Pew Research Center.
  • The Performance: When the actor sees a post on the screen, the AI "thinks" about it based on their character script and decides what to do (like, follow, read, or ignore). It's not a rigid rule; it's a natural reaction.
  • The Experiment: The researchers created 1,120 of these actors. They grouped them into 14 different "character types" (personas). For each character type, they created multiple copies (replicas).
    • The Twist: All copies of "Character A" act exactly the same way. However, the researchers changed just one visible detail for some of them. Maybe one copy says they are 25, while another says they are 55. Maybe one says they live in New York, another in Texas.

This setup allows the researchers to ask a powerful question: "If the exact same person (with the exact same behavior) changed their age or location, would the library show them different books?"

The Case Study: The 2024 Election

The team deployed these 1,120 AI actors on the social media platform X (formerly Twitter) right after the 2024 US Presidential Election. They watched how the platform's "For You" feed (the algorithmic one) compared to the "Following" feed (the chronological one).

Here is what they found:

1. The Algorithm Loves Drama and Right-Leaning Views
Compared to the standard chronological list, the "For You" feed showed significantly more:

  • Toxic content: Mean or angry posts.
  • Polarizing content: Posts that pit groups against each other.
  • Political content: Specifically, content leaning to the right.
  • Surprise: It did not amplify left-leaning content significantly. In fact, for users who leaned right, the algorithm actively suppressed left-leaning content.

2. The "Echo Chamber" is Asymmetric
The algorithm didn't just show everyone more right-wing stuff. It treated different users differently:

  • Left-leaning users saw a massive spike in toxic content (+80% more than the baseline).
  • Right-leaning users saw a spike in right-leaning content, but their toxic content didn't change much.
  • Result: Right-leaning users were put in a one-sided bubble where opposing views were hidden, while left-leaning users were flooded with anger and toxicity.

3. Demographics Matter, But It's Complicated
The researchers tested if changing a user's visible age, gender, or location changed what they saw.

  • The Average Result: If you look at all users together, changing these details didn't seem to matter much. The average effect was near zero.
  • The Real Story: When they looked at specific character types, the results were wild. Changing a signal (like age) might make one character see more toxic content, but make a different character see less.
  • The Metaphor: Imagine a thermostat that doesn't just turn the heat up or down based on the room temperature. Instead, it turns the heat up for some people and down for others based on who they are. If you just measure the "average" temperature of the whole house, you might think the thermostat isn't working at all. But if you look at each room, you see it's doing very specific, different things.

Why This Matters

The paper concludes that we can't just look at "average" results when auditing algorithms. If we average out the data, we miss the fact that the algorithm treats different groups of people in completely different ways.

By using these AI "actors," researchers can now run controlled experiments to see exactly how an algorithm reacts to specific traits, without needing to recruit thousands of real humans or rely on dumb, scripted bots. It's a new tool to understand the invisible rules that shape what we see online.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →