← Latest papers
💬 NLP

Language Models Should be Used to Surface the Unwritten Code of Science and Society

This paper proposes a framework that leverages large language models to uncover and critique society's "unwritten code" by analyzing their self-generated heuristics, demonstrating through a peer review case study how LLMs can reveal implicit biases—such as the unspoken preference for storytelling over pure rigor—that human experts apply but rarely articulate.

Original authors: Honglin Bao, Siyang Wu, Jiwoong Choi, Yingrong Mao, James A. Evans

Published 2026-01-28
📖 5 min read🧠 Deep dive

Original authors: Honglin Bao, Siyang Wu, Jiwoong Choi, Yingrong Mao, James A. Evans

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Finding the "Secret Rules"

Imagine you are trying to figure out why a specific song became a massive hit while another equally good song was ignored. You ask the music critics, "Why did this one win?" They give you a standard answer: "It had better lyrics and a catchy melody."

But deep down, you suspect there's a secret rule they aren't saying out loud. Maybe the winner just had a better story, or the critics liked the artist's background, or the song fit a current trend. These are the "unwritten codes"—the invisible rules that actually drive decisions, even though people are too polite or too used to them to admit it.

This paper argues that Large Language Models (LLMs)—the smart AI chatbots we use today—are actually perfect tools to find these secret rules. Instead of just trying to fix the AI's biases, the authors suggest we should use those biases as a flashlight to shine a light on the hidden, unwritten rules of human society.

The Experiment: The "Peer Review" Detective Story

To prove this, the researchers looked at scientific peer review. This is the process where scientists submit their papers to conferences, and other scientists (reviewers) decide if they are good enough to be published.

The Setup:

  1. The Data: They gathered thousands of pairs of scientific papers. In each pair, one paper got a high score (a "pass") and the other got a low score (a "fail").
  2. The AI Detective: They asked an AI to look at these pairs and guess: "Why did the high-scoring paper win?"

The Two-Step Discovery Process

The researchers didn't just ask the AI once. They used a clever two-step method to peel back the layers of the AI's thinking:

Step 1: The "Polite" Answer (The Prior)
First, they asked the AI what makes a paper good in general, without showing it any specific examples.

  • The Result: The AI gave very standard, "textbook" answers. It said things like, "The winner had better math," "The methods were more rigorous," or "The theory was stronger."
  • The Analogy: This is like a student giving the answer they think the teacher wants to hear. It's the "official" rulebook.

Step 2: The "Real" Answer (The Posterior)
Next, they showed the AI the actual pairs of papers (the winners and losers) and asked it to explain the difference specifically for those cases. They forced the AI to keep digging deeper until it could explain almost every single win.

  • The Result: The AI's answers changed! It started saying things like, "The winner told a better story," "It connected its work to other famous studies," or "It framed the problem in a way that felt important."
  • The Analogy: This is the student whispering, "Okay, but really, the teacher liked the one with the cool story and the flashy presentation, even if the math was the same."

The Shocking Discovery

The researchers found a huge gap between what humans say they value and what they actually reward.

  • What Humans Say: In their written reviews, human scientists mostly talked about the "textbook" stuff (rigor, methods). They rarely mentioned storytelling or how a paper was framed.
  • What Humans Actually Do: Despite writing about "rigor," the papers that won were the ones with better stories and connections.
  • The AI's Role: The AI started with the same "textbook" bias as the humans. But when forced to look at the real results, the AI updated its thinking. It realized, "Oh, I see. The real reason these papers won wasn't just the math; it was the story."

The Metaphor:
Imagine a judge at a talent show.

  • The Public Statement: The judge says, "I am looking for technical skill and pitch-perfect singing."
  • The Reality: The judge actually gives the trophy to the singer who tells the most emotional backstory and connects with the audience.
  • The AI's Job: The AI acts like a super-observant assistant who watches the judge's actual choices and says, "Hey, you keep giving trophies to the storytellers, even though you say you only care about pitch. Here is the secret rule you are following."

Why This Matters

The authors aren't saying we should let AI replace human judges. Instead, they want to use AI as a diagnostic tool.

  • Making the Invisible Visible: Just like this study revealed that "storytelling" is the hidden rule in science, this method could reveal hidden biases in hiring (e.g., "We say we want skills, but we actually prefer candidates who went to specific colleges") or college admissions.
  • The Goal: Once we know the secret rules, we can have an honest conversation about them. Are these rules fair? Do they help the right people? If we don't know the rules exist, we can't change them.

Summary

The paper claims that AI models have absorbed the hidden, unwritten rules of human culture. By forcing these AI models to explain their decisions over and over again, we can force them to "spill the beans" on what really matters in society—revealing the gap between our polite, official rules and the messy, real reasons we make choices. This helps us see the "hidden curriculum" of society so we can decide if we want to keep it or change it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →