← Latest papers
📊 statistics

TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design

This paper introduces TASTE, a multi-dimensional preference dataset annotated by professional designers across nine graphic design criteria, and demonstrates that while current AI evaluators struggle to match human consensus, a specialized model trained on this dataset significantly improves alignment with expert judgments.

Original authors: Haonan Zhu, Elad Hirsch, Alexandria Minetti, Allison Nulty, Purvanshi Mehta

Published 2026-05-21
📖 5 min read🧠 Deep dive

Original authors: Haonan Zhu, Elad Hirsch, Alexandria Minetti, Allison Nulty, Purvanshi Mehta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a boss trying to hire a graphic designer. You have two candidates, and you ask your team of five expert designers to pick the better one.

In the world of Artificial Intelligence (AI), we've been training computers to make images using a very simple rule: "Which photo looks more real?" or "Which picture is prettier?" This works great for making photos of cats or sunsets. But when it comes to graphic design (like posters, ads, or app layouts), "pretty" isn't enough. A poster might look beautiful but have the wrong font, the text might be misspelled, or the colors might clash with the brand.

This paper introduces TASTE, a new tool to teach AI how to be a real graphic designer, not just a pretty picture generator.

Here is the breakdown of what they did, using some everyday analogies:

1. The Problem: The "One-Size-Fits-All" Judge

Imagine you are judging a cooking contest.

  • Old AI Judges: They only ask, "Does this dish look delicious?" If the food looks good, they give it a high score, even if it's salty or the wrong temperature.
  • The Reality: Real chefs (human designers) judge food on many different things: Is it seasoned right? Is the plating neat? Did the chef follow the recipe? Did they use the right ingredients?

The authors found that current AI models are being trained on "food that looks delicious" (photo-style data), but they need to be trained on "food that follows the recipe and tastes right" (graphic design data). A single "thumbs up" label isn't enough because it hides why a design failed.

2. The Solution: The "TASTE" Dataset

The team created a dataset called TASTE (Typography, Aesthetics, Spatial, Tone, Etc.). Think of this as a massive, detailed report card.

  • The Players: They hired 10 professional human designers.
  • The Contest: They asked four different AI image generators to create designs based on specific prompts (like "Make a flyer for a coffee shop").
  • The Grading: Instead of just saying "I like this one," the designers graded the AI on nine specific categories:
    • Typography: Is the font choice and spacing good?
    • Visual Hierarchy: Can you tell what's the most important thing to look at?
    • Color Harmony: Do the colors work well together?
    • Spatial Accuracy: Is the layout where it's supposed to be?
    • Hallucinations: Did the AI invent things that weren't asked for (like a dog in a coffee shop ad)?

They did this for 1,600 different ratings per category. This gives a very clear picture of what human designers actually care about.

3. The "Truth Test": Are the Humans Agreeing?

Before they could use this data to train AI, they had to make sure the human designers weren't just guessing randomly. They used three statistical "lie detector" tests:

  1. Do they agree on the ranking? (Like asking 5 people to rank their favorite ice cream flavors).
  2. Is there a clear winner? (Do most people agree on the top choice, or is it a total tie?).
  3. Are there logical loops? (If Person A likes X over Y, and Y over Z, do they also like X over Z? Or do they get confused?).

The Result: The humans agreed much more than random guessers, but less than people agreeing on "is this photo blurry?" This tells us that graphic design is a mix of objective rules (like spelling) and subjective taste (like color vibes). It's a "sweet spot" between following a recipe and being an artist.

4. The Reality Check: Can Current AI Judges Do This?

The authors tested the smartest AI judges available today (the "referees") to see if they could predict what the human designers would pick.

  • The Score: The best AI judges got about 54% accuracy.
  • The Analogy: This is like flipping a coin. If you flip a coin to guess who the human designers will pick, you'd get it right 50% of the time. The AI judges are only slightly better than a coin flip.
  • The Problem: Bigger AI models didn't get much better. It seems like the current AI judges are great at spotting "pretty photos" but terrible at spotting "good design." They get confused by text, layout, and specific instructions.

5. The New "Coach": A Small, Specialized Model

Since the big, general AI judges failed, the authors built a small, specialized "coach" (a small AI model) trained specifically on the TASTE data.

  • The Trick: Instead of just looking at one image at a time, this coach looks at two images side-by-side and asks, "What is the difference between these two that makes one better?"
  • The Result: This new coach reached 61% accuracy.
  • The Ceiling: The authors calculated that even a perfect AI could only reach about 74% accuracy because sometimes even humans disagree. So, this new coach has closed about half the gap between "random guessing" and "human expert."

Summary

This paper says: "We can't just teach AI to make pretty pictures. We need to teach it the specific rules of design (fonts, layout, colors) using a new dataset called TASTE. Current AI judges are failing at this, but a small, specialized model trained on our new data is starting to learn the ropes. We are now giving this data to the world so others can build better design tools."

What the paper does NOT claim:

  • It does not claim that AI can now replace human designers.
  • It does not claim that this solves all design problems (like accessibility or brand consistency, which they admit are missing).
  • It does not claim that bigger models will automatically fix the problem; in fact, they found that just making models bigger didn't help much.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →