← Latest papers
💻 computer science

MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data

To overcome the performance degradation of current models in multi-reference image generation caused by a lack of structured long-context data, this paper introduces MacroData, a large-scale dataset of 400K samples, and MacroBench, a standardized evaluation benchmark, which together enable substantial improvements in generating coherent images from multiple visual references.

Original authors: Zhekai Chen, Yuqing Wang, Manyuan Zhang, Xihui Liu

Published 2026-03-27
📖 4 min read☕ Coffee break read

Original authors: Zhekai Chen, Yuqing Wang, Manyuan Zhang, Xihui Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a very talented artist how to paint.

The Problem: The "One-Reference" Bottleneck
Right now, most AI image generators are like students who have only ever practiced with one reference photo at a time.

  • If you show them a photo of a cat and say, "Draw this cat in a hat," they do great.
  • But if you say, "Here are 8 photos: a cat, a hat, a red scarf, a sunny park, a bicycle, a picnic basket, a dog, and a tree. Now, draw a scene where all these things interact naturally," the AI gets confused. It might forget the hat, mix up the dog's face with the cat's, or drop the bicycle entirely.

Current AI models struggle when you give them too many visual clues at once. They suffer from "visual amnesia."

The Solution: "MacroData" (The Ultimate Art Class)
The authors of this paper realized the problem wasn't the AI's intelligence; it was the textbook they were studying from. Existing textbooks only had exercises with 1 or 2 pictures.

So, they created MacroData, a massive new "textbook" containing 400,000 examples.

  • The Scale: Instead of showing the AI 1 or 2 pictures, this dataset forces the AI to practice with up to 10 different images in a single prompt.
  • The Analogy: Think of it like a cooking class. Old textbooks only taught you how to make a sandwich (1 ingredient). This new textbook teaches you how to make a complex banquet where you have to juggle 10 different ingredients simultaneously, ensuring the flavors blend perfectly without one overpowering the others.

The Four "Courses" in the Class
To make sure the AI learns to handle any situation, they organized the 400,000 examples into four distinct "courses":

  1. Customization (The "Mash-Up" Course):

    • The Task: Take a person from Photo A, a shirt from Photo B, and a background from Photo C, and combine them into one perfect image.
    • The Metaphor: Like a DJ mixing 10 different songs into one seamless track without the beat dropping.
  2. Illustration (The "Storyteller" Course):

    • The Task: You give the AI a long story with pictures scattered throughout. It needs to draw the next picture that fits the story perfectly.
    • The Metaphor: Like a comic book artist who has read the first 10 pages of a graphic novel and needs to draw page 11 so it matches the characters' expressions and the setting exactly.
  3. Spatial (The "3D Detective" Course):

    • The Task: Show the AI a cube from the front, top, and right side. Ask it to draw what the cube looks like from the back-left corner.
    • The Metaphor: Like a real estate agent showing you photos of a house from the street, the backyard, and the roof, and asking you to draw a photo of the house from the side alley you've never seen.
  4. Temporal (The "Time Traveler" Course):

    • The Task: Show the AI a sequence of 8 video frames of a ball bouncing. Ask it to draw the 9th frame (where the ball will be next).
    • The Metaphor: Like watching a movie and pausing it, then predicting exactly what the next frame will look like based on the momentum of the previous ones.

The Exam: "MacroBench"
You can't just say, "We made a better textbook." You have to prove the students passed the test.
The authors also built MacroBench, a standardized exam.

  • The Twist: Most exams only test you on 1 or 2 pictures. This exam gets harder as you go. It tests the AI with 1 picture, then 5, then 10.
  • The Result: When the AI was trained on the new "MacroData" textbook, it didn't just get slightly better; it became a master at handling complex, multi-image tasks, closing the gap between open-source AI and the expensive, "closed-source" giants (like the ones used by big tech companies).

Why This Matters
Before this paper, if you wanted to create a complex image with many specific elements, you had to use expensive, closed tools or do it manually.

  • The Old Way: Trying to build a house with a hammer that only works on one nail at a time.
  • The New Way: This paper gives the AI a power drill that can handle 10 nails at once, perfectly aligned.

In a Nutshell
The paper says: "AI is smart, but it's been studying the wrong books. We wrote a new, massive library of books that teach AI how to juggle 10 visual ideas at once. Now, the AI can create complex, coherent images that were previously impossible."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →