← Latest papers
💻 computer science

Email in the Era of LLMs

This paper introduces the HR Simulator, a communication game demonstrating that while large language models (LLMs) often outperform humans in formal email scenarios and exhibit more homogenous, tactful preferences, human-LLM collaboration significantly enhances success rates, revealing both the efficacy of LLMs in structured communication and current limitations in mimicking human-like low-empathy, informal styles.

Original authors: Dang Nguyen, Harvey Yiyun Fu, Peter West, Chenhao Tan, Ari Holtzman

Published 2026-03-24
📖 5 min read🧠 Deep dive

Original authors: Dang Nguyen, Harvey Yiyun Fu, Peter West, Chenhao Tan, Ari Holtzman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're playing a high-stakes video game called "HR Simulator." In this game, you aren't fighting dragons or racing cars; you are the new Human Resources officer at a fictional tech company called "NeuroGrid." Your only weapon? Email.

Every level of the game presents a tricky social situation: a new employee wants a private office they aren't eligible for, two coworkers are fighting over a project, or a senior engineer is secretly unhappy but won't say why. Your goal is to write an email that solves the problem without starting a war, making someone cry, or getting yourself fired.

The twist? In this future world, your emails aren't just read by humans. They are read, judged, and sometimes even written by AI (Large Language Models).

Here is what the researchers discovered by playing this game with 600+ emails:

1. The AI Judges are Getting "Groupthinky"

Imagine a panel of judges tasting a soup.

  • Small AI models are like a group of friends who all have different opinions. One says it needs salt, another says it needs pepper. They can't agree on what "good" tastes like.
  • Big, powerful AI models are like a group of food critics who have read the same cookbook. They all agree: "Good soup must be perfectly seasoned and warm."

As AI gets smarter, they all start agreeing on what a "good email" looks like. They are converging on a single, very specific style. The problem? That style isn't always what humans prefer.

2. Humans vs. AI: The "Polite Robot" Problem

When humans wrote emails alone, they were often too blunt, too casual, or just "off."
When AI wrote emails alone, they were too perfect. They were like a robot trying to be a human: overly formal, excessively empathetic, and sometimes a bit robotic in their kindness.

The Result: When an AI judge graded the emails, the AI-written ones usually won. Humans lost. It's like a human trying to sing a song against a perfectly tuned synthesizer; the machine hits every note, but the human has "soul" that the machine misses.

3. The Magic Combo: Human + AI (The "Cyborg" Strategy)

Here is the best news: The best emails were written by a Human and an AI working together.

Think of it like this:

  • The Human is the Architect. They know the goal: "I need to tell Sam he can't have a private office, but I need to keep him excited about the job." They know the nuance of the situation.
  • The AI is the Interior Designer. They take the human's rough draft and polish it. They add the right amount of "warmth," fix the grammar, and make sure the tone sounds professional.

When you combine them, the success rate skyrocketed. In some difficult scenarios, a human alone had a 40% chance of winning. An AI alone had a 50% chance. But Human + AI together had nearly a 100% chance.

4. The "Tact" Discovery: Bigger AI = More Subtle

The researchers found something fascinating called "Emergent Tact."

  • Smaller AI models were like blunt instruments. They preferred emails that were direct, maybe a bit aggressive, or tried to "buy" the other person's agreement with perks (like "You can't have an office, but here's a noise-canceling headset!").
  • Larger, smarter AI models were like master diplomats. They preferred subtlety. They knew that saying "You can't have an office" is bad, but framing it as "We value collaboration in our open space" is good.

As AI gets smarter, it learns that being tactful is the key to winning.

5. The One Thing AI Can't Do: The "Grumpy Text"

There is one corner of the email world where AI fails miserably: The "Low Empathy, Low Formality" zone.

Imagine you are so frustrated with a company's bureaucracy that you want to write a short, angry, rude email just to get their attention.

  • Humans can do this easily. We can be curt, dismissive, and angry when we need to be.
  • AI cannot. Even when asked to be rude, the AI keeps trying to be polite and helpful. It's like asking a golden retriever to bite someone; it just wags its tail and offers a treat.

This is important because sometimes, in real life, you need to be blunt and unemotional to get a point across. AI can't do that alone. This is another reason why we need humans in the loop—to provide that "grumpy" edge when necessary.

The Big Picture: What Does This Mean for Us?

The paper suggests a future where:

  1. If you try to compete with AI alone, you will likely lose. AI is getting too good at the "standard" rules of polite communication.
  2. If you team up with AI, you become unstoppable. You bring the human intuition and the "grumpiness" when needed; the AI brings the polish and the diplomatic finesse.

The Verdict: The future of email isn't "Human vs. Machine." It's "Human + Machine." We shouldn't try to out-write the robots; we should let them be our editors while we remain the architects of our messages.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →