← Latest papers
💻 computer science

MS-CustomNet: Controllable Multi-Subject Customization with Hierarchical Relational Semantics

MS-CustomNet is a novel framework that enables zero-shot, fine-grained control over multi-subject image generation by allowing users to explicitly define hierarchical arrangements and spatial relationships while preserving individual subject identities, supported by a new MSI dataset and demonstrating superior performance in identity preservation and positional control.

Original authors: Pengxiang Cai, Mengyang Li

Published 2026-03-24
📖 5 min read🧠 Deep dive

Original authors: Pengxiang Cai, Mengyang Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a magical photo editor that can turn your text descriptions into real pictures. You can tell it, "Draw a cat on a mat," and it does. But what if you want something more specific? What if you want to say, "Put this specific cat (the one from your phone) sitting inside this specific bowl (the one from your kitchen), while that specific dog (your neighbor's) watches from behind the bowl"?

Current AI tools are like talented but slightly chaotic painters. They are great at painting a cat or a bowl, but if you ask them to put three specific things together in a very precise way, they often get confused. They might mix the cat's face with the dog's ears, or they might put the bowl floating in the sky instead of on the table. They struggle to understand the "rules" of how objects stack, hide behind each other, or interact.

Enter MS-CustomNet. Think of it as a master stage director for a play, rather than just a painter.

Here is how it works, broken down into simple concepts:

1. The Problem: The "Concept Bleed"

Imagine you have a box of LEGO bricks. You want to build a castle, a spaceship, and a robot. If you just throw all the bricks in a pile and say "Build me a scene," the AI might accidentally glue the robot's arm onto the spaceship's wing. This is called "concept bleeding." Existing tools often fail to keep the characters distinct or to follow your instructions on who is standing in front of whom.

2. The Solution: The "Director's Script" (MS-CustomNet)

MS-CustomNet is a new system designed to be the ultimate director. It doesn't just guess; it follows a script you give it.

  • The Cast (Identity Preservation): You show the AI photos of your specific "actors" (a specific backpack, a specific cup, a specific dog). The system has a special memory trick that ensures the backpack stays a backpack and the dog stays a dog, even when they are moved to a new scene. It's like having a costume department that never loses the actors' unique features.
  • The Stage Directions (Spatial Control): This is the magic part. You don't just say "put them together." You give the AI a layout map. Think of this like a floor plan for a theater. You tell the AI: "The cup goes inside the bowl," or "The book goes on top of the stack." The AI reads this map and builds the scene exactly to those instructions, respecting who is in front of whom and who is hiding behind whom.

3. The Training Ground: The "MSI Dataset"

To teach this AI to be so good at following directions, the researchers had to build a new school. They created a massive dataset called MSI (Multi-Subject Interaction).

  • Imagine taking thousands of photos from a library (the COCO dataset).
  • They filtered out the boring ones and kept only the complex scenes with multiple objects interacting (like a dog playing with a ball near a person).
  • They then created "training exercises" where the AI had to learn how to take a photo of a dog, a ball, and a person, and rearrange them according to a new set of rules. This is like a student practicing with flashcards until they can solve any puzzle.

4. The Learning Strategy: "Baby Steps" (Curriculum Learning)

You wouldn't ask a toddler to run a marathon on day one. You start with walking, then jogging, then running.
The researchers used a similar strategy called Curriculum Learning.

  • Phase 1: They taught the AI to handle just two objects (e.g., a cup and a saucer).
  • Phase 2: Once the AI got good at that, they added a third object.
  • Phase 3: Finally, they let it handle up to five objects at once.
    This prevented the AI from getting overwhelmed and confused, allowing it to learn complex interactions step-by-step.

5. The Result: A Perfectly Staged Scene

When you use MS-CustomNet, you get a result that feels like a movie set:

  • The Identity: The objects look exactly like the photos you gave it.
  • The Position: The objects are exactly where you told them to be.
  • The Interaction: If you said "the cup is inside the bowl," the AI knows to draw the rim of the bowl covering the top of the cup. It understands depth and layering.

In summary:
If current AI image generators are like a magician who pulls random rabbits out of a hat, MS-CustomNet is like a chef who follows your exact recipe. It takes the specific ingredients you provide, arranges them exactly how you want them on the plate, and ensures the final dish looks delicious and coherent. It gives you the power to be the director of your own visual story, with total control over who stands where and how they interact.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →