← Latest papers
💬 NLP

HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing

This paper introduces HoWToBench, a large-scale Chinese writing benchmark, and proposes Tree-of-Writing (ToW), a tree-structured evaluation framework that explicitly models sub-feature aggregation to achieve high correlation with human judgments and robustness against textual disturbances, thereby addressing the limitations of existing metrics in assessing LLMs' human-level writing capabilities.

Original authors: Andrew Zhuoer Feng, Cunxiang Wang, Yu Luo, Lin Fan, Yilin Zhou, Zikang Wang, Xiaotao Gu, Jie Tang, Hongning Wang, Minlie Huang

Published 2026-04-22
📖 5 min read🧠 Deep dive

Original authors: Andrew Zhuoer Feng, Cunxiang Wang, Yu Luo, Lin Fan, Yilin Zhou, Zikang Wang, Xiaotao Gu, Jie Tang, Hongning Wang, Minlie Huang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Robot Judge" is Getting Confused

Imagine you are a famous chef. You've just cooked a complex, 10-course meal. You want to know if it's good.

In the past, to judge AI writing, we used two main methods:

  1. The "Copy-Paste" Judge: We compared the AI's writing to a human's writing word-for-word. If they didn't match perfectly, the AI failed. This is like judging a jazz musician by how well they can play a song exactly like a recording. It misses the creativity.
  2. The "One-Shot" Robot Judge: We asked another AI to read the writing and give it a score (e.g., "8 out of 10"). But here's the catch: The robot judge is bad at math. When it tries to combine scores for "grammar," "creativity," and "logic," it often just averages them out or gets confused. It might give a high score to a boring essay just because it was long, or a low score to a brilliant story just because it had one typo.

The researchers call this "Negotiation Inconsistency." It's like asking a committee of robots to decide the winner of a talent show, but every time you ask them, they change their minds because they can't agree on how to weigh the different talents.

The Solution: The "Tree of Writing" (ToW)

The authors built a new way to judge writing called Tree-of-Writing (ToW).

The Metaphor: The Tree vs. The Smoothie

  • Old Way (The Smoothie): You throw all the ingredients (grammar, plot, emotion, formatting) into a blender. You get a smoothie (a single score). You can't taste the strawberry separately from the banana. If the blender is broken, the whole drink tastes bad.
  • New Way (The Tree): Imagine a tree.
    • The Root is the final score.
    • The Main Branches are big categories: Content (the story), Format (the structure), and Impression (the "vibe").
    • The Leaves are tiny details: "Did the opening grab you?" "Is the logic sound?" "Are the paragraphs the right size?"

How it works:
Instead of one robot guessing the whole score, the system sends different "Expert Agents" to specific leaves.

  • One agent checks if the logic makes sense.
  • Another checks if the formatting (like headers and lists) is correct.
  • A third checks the emotional depth.

Crucially, a "Negotiator" (a smart AI) decides how much each leaf matters before the scoring starts. For a poem, "Emotion" gets a heavy weight. For a legal contract, "Logic" gets the heavy weight. This prevents the robot from getting confused about what is important.

The New Test: HoWToBench

To test this new judging system, they built a massive test called HoWToBench.

The Metaphor: The "Master Chef's Kitchen"
Most AI tests are like asking a robot to fill in a blank in a sentence. It's easy.
HoWToBench is like giving the robot a blank canvas and saying: "Write a mystery novel," or "Write a speech for a mayor," or "Write a contract for a space hotel."

  • 12 Genres: They tested everything from poetry and fiction to official government documents and business plans.
  • 3 Difficulty Levels:
    1. Completion: "Here is a story, finish the middle." (Easy)
    2. Guide: "Here is an outline, write the story." (Medium)
    3. Open: "Here is a topic, write whatever you want." (Hard)

They used real human experts (writers, journalists, editors) to grade the AI's work, creating a "Gold Standard" to see if the Tree-of-Writing system was actually right.

The Surprising Discoveries

When they ran the tests, they found some things that surprised everyone:

  1. Longer is NOT Better:

    • The Myth: "If the AI writes a huge essay, it must be smart."
    • The Reality: The study found that longer inputs and outputs often got lower scores. If an AI just piles on words to look smart, it actually hurts its score. Humans prefer concise, punchy writing over rambling nonsense.
    • Analogy: It's like a comedian who tells a joke. If they tell the punchline in 5 seconds, it's funny. If they talk for 20 minutes before getting to the point, it's not funny.
  2. The "Fragile" Judges:

    • They tested if the judges were robust. They took a good AI story and did weird things to it: deleted a paragraph, repeated a paragraph, or changed the genre.
    • Result: Old judging methods (like simple word-matching or standard AI judges) got confused and gave wildly different scores.
    • The Tree: The Tree-of-Writing stayed calm. It recognized that the story was still good even if a paragraph was missing, or that the repetition was bad. It was robust.
  3. The "Mimicry" Trap:

    • Many AIs are great at following instructions when they have a lot of help (like a recipe). But when you take away the recipe (Open Writing), they struggle to be truly creative. They are good at mimicking humans, but not yet at being human-level writers.

Why This Matters

This paper is a big step forward because it stops treating writing like a math problem (counting words) and starts treating it like art.

By using a Tree structure, they made the judging process transparent. We can now see why an AI got a bad score (e.g., "It failed on Logic, but did great on Emotion"). This helps developers fix their models and helps us trust AI when it writes stories, contracts, or speeches.

In short: They built a smarter, more human-like way to grade AI writing, proving that quality isn't about how long you write, but how well you write.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →