← Latest papers
💻 computer science

MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation

This paper introduces MultiBanana, a comprehensive open benchmark designed to evaluate and advance multi-reference text-to-image generation models by assessing their capabilities across diverse challenges such as varying reference counts, domain and scale mismatches, rare concepts, and multilingual inputs.

Original authors: Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

Published 2026-03-27
📖 5 min read🧠 Deep dive

Original authors: Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a master chef in a kitchen. Until recently, your job was simple: take one photo of a strawberry and one recipe, and bake a cake that looks exactly like that strawberry. You were getting really good at this.

But now, the customers are getting fancy. They don't just want a strawberry cake. They want a cake that uses:

  • The strawberry from Photo A.
  • The chocolate frosting from Photo B.
  • The blueberry from Photo C.
  • The sprinkles from Photo D.
  • The shape of the plate from Photo E.
  • The lighting from Photo F.
  • The style of the tablecloth from Photo G.
  • And the text on the menu from Photo H.

And they want it all to look like it belongs in the same room, not like a messy collage.

This is the problem the paper "MultiBanana" is trying to solve.

The Problem: The "Too Many Ingredients" Crisis

Current AI image generators (like the ones making pictures for ads or social media) are great at using one or two reference photos. But when you give them a "shopping list" of 8 different images to combine, they start to panic. They might forget the blueberries, mix up the chocolate with the sprinkles, or make the cake look like it's floating in mid-air.

The researchers realized that the old tests (benchmarks) were like testing a chef with only one ingredient. They didn't tell us if the chef could actually handle a complex banquet.

The Solution: Introducing "MultiBanana"

The team from the University of Tokyo and Google DeepMind created a new, super-hard test called MultiBanana.

Think of MultiBanana as a "Culinary Olympics" for AI chefs. Instead of asking them to bake a simple cookie, they throw a massive challenge at them:

  1. The Quantity Challenge: Can you combine 3, 4, or even 8 different images at once?
  2. The Style Clash: Can you mix a realistic photo of a dog with a cartoon cat and a watercolor painting of a background, and make them look like they belong together?
  3. The Size Problem: Can you take a tiny ant from one photo and a giant mountain from another, and make them fit in the same scene without the ant looking like a toy?
  4. The Rare Ingredient: Can you bake a cake using a picture of a "red banana" (which doesn't really exist in nature)?
  5. The Language Barrier: Can you read a sign in Japanese, a menu in Chinese, and a recipe in English, and put them all on the same picture correctly?

How They Tested the Chefs

They didn't just ask humans to look at the pictures (that takes too long and costs too much). Instead, they used other super-smart AIs (like GPT-5 and Gemini) to act as judges. These judges looked at the results and gave them a score out of 10 based on five things:

  • Did they follow the recipe? (Did they use all the ingredients?)
  • Did the ingredients look right? (Did the red banana look like the red banana in the photo?)
  • Does it look natural? (Is the dog sitting on the ground, or floating?)
  • Is the lighting right?
  • Is it pretty?

What They Found: The "Over-Thinker" vs. The "Under-Thinker"

When they ran the test on the top AI models, they found two very different types of failures:

1. The "Over-Thinker" (Closed-Source Models like Nano Banana & GPT-Image-1)
These models are like chefs who try so hard to use every single ingredient that they ruin the dish.

  • What they do: They successfully grab all 8 items.
  • The flaw: Because they are trying so hard to copy every tiny detail, the final picture looks weird. The characters might be squished together, the lighting is confusing, or the scene looks "mushy" and distorted. They are faithful to the ingredients but bad at the composition.

2. The "Under-Thinker" (Open-Source Models like Qwen & DreamOmni2)
These models are like chefs who get overwhelmed by the list and just give up on some ingredients.

  • What they do: They make a very clean, beautiful picture.
  • The flaw: They forget half the order. If you asked for 8 items, they might only put 3 in the picture. They ignore the instructions to keep the image looking "safe" and pretty.

Why This Matters

This paper is a wake-up call. It tells us that while AI is getting amazing at making one thing look real, it's still struggling to be a director who can manage a whole cast of characters from different movies and put them in one scene.

MultiBanana is now the new standard. It's a public playground where anyone can test their AI to see if it can actually handle the chaos of real-world creativity. It pushes developers to stop just making pretty pictures and start teaching their AI how to follow complex, multi-step instructions without losing its mind.

In short: MultiBanana is the ultimate stress test to see if AI can finally stop being a "one-trick pony" and become a true "master chef" of the visual world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →