← Latest papers
💻 computer science

Open-Domain Safety Policy Construction

This paper introduces Deep Policy Research (DPR), a minimal agentic system that autonomously drafts comprehensive content moderation policies from seed domain information by iteratively searching the web and distilling rules, demonstrating performance that surpasses existing baselines and rivals expert-written policies across diverse benchmarks.

Original authors: Di Wu, Siyue Liu, Zixiang Ji, Ya-Liang Chang, Zhe-Yu Liu, Andrew Pleffer, Kai-Wei Chang

Published 2026-04-03
📖 4 min read☕ Coffee break read

Original authors: Di Wu, Siyue Liu, Zixiang Ji, Ya-Liang Chang, Zhe-Yu Liu, Andrew Pleffer, Kai-Wei Chang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you run a massive, bustling digital town square. In this square, people post ads, share stories, and chat. But sometimes, people say mean things, show scary images, or try to trick others. To keep the town safe, you need a Rulebook (a safety policy) that tells your security guards exactly what is allowed and what isn't.

Usually, writing this Rulebook is a huge, expensive headache. You have to hire experts who know every possible "edge case" (like, "Is a cartoon knife okay? What about a joke that sounds like a threat?"). They have to write thousands of rules, update them constantly, and argue over the wording.

This paper introduces a new robot assistant called "Deep Policy Research" (DPR) that writes this Rulebook for you.

Here is how it works, using a simple analogy:

🕵️‍♂️ The Detective vs. The Librarian

Think of the old way of making rules as hiring a Librarian who sits in a quiet room, tries to remember every rule they've ever read, and writes a list based on their memory. Sometimes they forget the tricky details.

DPR is more like a Super-Detective.
You give the Detective a one-sentence clue, like: "We need rules about 'Violence'."
Instead of guessing, the Detective:

  1. Goes to the Library (The Internet): It doesn't just guess; it actually searches the web. It looks at news sites, government documents, and forums to see how real people define violence.
  2. Asks Specific Questions: It asks things like, "What counts as violence in a cartoon?" or "How do we handle jokes about fighting?"
  3. Takes Notes: It reads the answers and turns them into clear, simple rules.
  4. Organizes the Filing Cabinet: It doesn't just dump a pile of notes on your desk. It sorts them into neat folders (like "Physical Violence," "Verbal Threats," "Gore") so the security guards can find the right rule quickly.

🏗️ How the Robot Builds the Rulebook

The paper describes a three-step loop that the robot repeats a few times:

  1. The Search: The robot looks at what it has so far and asks, "What am I missing?" It then searches the web for those missing pieces.
  2. The Translation: It finds messy, long articles on the web and translates them into short, punchy rules.
    • Web Article: "There are many nuances to hate speech involving identity tourism..."
    • Robot Rule: "Do not allow content where someone pretends to be part of a group they aren't, just to mock them."
  3. The Organization: It takes all these new rules and groups them together, giving each group a clear title so the final document is easy to read.

🧪 Did It Work?

The researchers tested this robot in two ways:

  1. The Text Test: They asked the robot to write rules for five tricky topics: Sex, Hate, Violence, Harassment, and Self-Harm. They then gave these robot-written rules to a standard AI "Security Guard" to see if it could spot bad content better.

    • Result: The robot-written rules made the Security Guard much smarter. It caught more bad stuff than when the guard was just given a vague definition or a few examples. In fact, the robot did almost as well as a team of human experts.
  2. The Ad Test: They tested it on a real-world advertising company (Taboola) that has to check thousands of ads with pictures and text.

    • Result: When they replaced a human-written section of their rulebook with the robot's version, the system worked almost just as well! The robot was especially good at catching tricky, visual tricks (like ads that use shocking images to sell products).

🌟 The Big Takeaway

The most exciting part isn't just that the robot is smart; it's that it's simple.

  • It doesn't need a super-complex brain or a million tools.
  • It just needs one tool (a web search) and a simple plan (Search -> Summarize -> Organize).

The Metaphor:
Imagine you are trying to build a fence around a garden.

  • The Old Way: You hire a master builder to design the fence from scratch. It takes weeks and costs a fortune.
  • The DPR Way: You give a smart robot a shovel and a map. The robot runs around the garden, checks the terrain, finds the best spots for fence posts, and builds a sturdy fence that fits the garden perfectly.

This paper shows that we don't need to be the experts to write the rules anymore. We just need to build a robot that knows how to research like an expert. This could save companies millions of dollars and make the internet safer much faster.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →