← Latest papers
💬 NLP

Small edits, large models: How Wikipedia advocacy shapes LLM values

This paper demonstrates that a small, coordinated group of Wikipedia volunteers can significantly shape the values and responses of large language models regarding specific topics like animal welfare, as their targeted edits become disproportionately influential in model training and behavior compared to unrelated content.

Original authors: Jasmine Brazilek, Maria Navas, Alexa Gnauck

Published 2026-06-25
📖 5 min read🧠 Deep dive

Original authors: Jasmine Brazilek, Maria Navas, Alexa Gnauck

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, endless library where every book ever written is stacked on shelves. Now, imagine that Artificial Intelligence (AI) is a super-smart student who wants to learn everything about the world. To do this, the student doesn't read just one book; they read every book in the library.

However, this student is very picky. They don't read all the books equally. They pay extra attention to the "Encyclopedia" section (Wikipedia) because it's written by experts and is very reliable. In fact, the student reads the Encyclopedia pages three to five times more often than the random news articles or blog posts scattered around the rest of the library.

This paper is about a small group of volunteers called Pro-Animal Wikipedians (PAW). They decided to use this "super-student's" study habits to their advantage. Here is the simple breakdown of what they did and what they found:

1. The Strategy: Editing the Encyclopedia

The PAW group didn't try to hack the AI or build a new computer. Instead, they did something very simple: they edited Wikipedia.

They went to 115 different Wikipedia pages (mostly about big fast-food chains, politicians, and animal laws) and added 125 specific, fact-checked paragraphs about animal welfare.

  • Example: On a page about a fast-food chain, they added a section detailing exactly how that company treats its chickens.
  • The Goal: They wanted to see if changing the "Encyclopedia" would change what the "super-student" (the AI) learned about those companies.

2. The Experiment: Did the Student Notice?

The researchers wanted to know: Did the AI actually learn from these specific edits, or did it just ignore them?

To find out, they used three different "magnifying glasses" (scientific tools) to look at how the AI (specifically models like Llama) thought about animal welfare.

  • Magnifying Glass #1 (The "Memory Test"): They asked the AI questions like, "What is this company's policy on animal welfare?" The AI looked back at its training data to find the answer. The researchers found that 68% of the time, the AI's answer was directly pulled from the specific paragraphs the PAW group had written. Without those edits, the AI wouldn't have known those facts.
  • Magnifying Glass #2 (The "What If" Test): They ran a simulation asking, "If we erased the PAW edits from the AI's memory, would the AI get worse at answering animal welfare questions?" The answer was a resounding yes. In fact, in every single test run, the top 10 most important pieces of information the AI used to answer animal welfare questions were all the edits made by the PAW group.
  • Magnifying Glass #3 (The "Specialist" Test): They trained two tiny AI models from scratch. One only read the PAW edits; the other only read random Wikipedia text.
    • The "PAW Model" became an expert at animal welfare but knew nothing about general company history.
    • The "Random Model" became an expert on general history but knew nothing about animal welfare.
    • The Lesson: The AI didn't just memorize the words; it learned the ideas and associations specifically from the text it was fed.

3. The Crucial Discovery: It's Specific, Not General

The most important finding is that the AI didn't get confused.

  • When asked about animal welfare, the AI relied heavily on the PAW edits.
  • When asked about unrelated things (like "How many stores does this company have?"), the AI ignored the PAW edits and went back to its general knowledge.

The Analogy: Imagine you are teaching a child about a specific tree. You add a new, bright red leaf to the tree in a picture book.

  • If you ask, "What color is the leaf?" the child points to your red leaf.
  • If you ask, "How tall is the tree?" the child ignores the leaf and looks at the trunk.
    The child learned that the red leaf is important only for questions about leaves, not for questions about the whole tree. The AI did the exact same thing.

4. Why This Matters (According to the Paper)

The paper concludes that you don't need millions of dollars or a team of engineers to influence AI. You just need a small, coordinated group of volunteers to edit Wikipedia.

Because AI models are trained on Wikipedia, and because they weigh Wikipedia very heavily, changing the Wikipedia page is like changing the source code of the AI's knowledge.

  • The Scale: The paper notes that even though the AI models they tested were smaller than the ones you might use today (like the ones in your phone), the effect is likely even stronger in the bigger models. This is because the bigger models have more "memory" to hold onto these specific facts, and the Wikipedia facts are reinforced by other news sources the AI also reads.
  • The Takeaway: A small group of people, by simply writing factual, well-sourced articles on Wikipedia, can measurably change how AI systems talk about animal welfare. It is a low-cost, high-impact way to shape the "conscience" of artificial intelligence.

In short: The paper proves that if you want an AI to care about a specific topic, you don't need to program the AI. You just need to write the encyclopedia entry that the AI reads.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →