← Latest papers
🤖 machine learning

What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models

This paper investigates the drivers of compositional generalization in visual generative models, identifying that training on continuous distributions and providing informative conditioning are key factors, and demonstrates that augmenting discrete models with an auxiliary continuous JEPA-based objective improves their ability to generate novel concept combinations.

Original authors: Karim Farid, Rajat Sahay, Yumna Ali Alnaggar, Simon Schrodi, Volker Fischer, Cordelia Schmid, Thomas Brox

Published 2026-04-28
📖 4 min read☕ Coffee break read

Original authors: Karim Farid, Rajat Sahay, Yumna Ali Alnaggar, Simon Schrodi, Volker Fischer, Cordelia Schmid, Thomas Brox

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a child how to play with LEGO bricks. You show them a red square brick, a blue triangle brick, and a yellow circle brick. Then, you show them a house made of red squares and a car made of blue triangles.

If the child has truly learned the "rules of the game" (compositional generalization), you should be able to hand them a blue square and a yellow circle, and they should be able to build something new without getting confused.

But if the child has only "memorized the pictures" (interpolation), they might look at the blue square and yellow circle and say, "I don't know what to do with these; they don't look like the house or the car I saw before!"

This paper investigates why some AI models are like the smart child who understands the rules, while others are like the child who just memorizes the pictures.

The Discovery: The "Smoothness" Secret

The researchers compared two main types of AI "brains":

  1. The Discrete Brain (The "Checkbox" Model): This model thinks in rigid categories. To it, colors and shapes are like a series of checkboxes. It sees "Red," "Blue," "Square," or "Circle." There is no middle ground. It’s like a person who can only communicate using a limited set of stickers.
  2. The Continuous Brain (The "Slider" Model): This model thinks in a spectrum. To it, colors and shapes are like sliding scales. It understands that "Red" is a specific point on a long, smooth rainbow.

The Finding: The researchers found that the Continuous Brain is much better at playing with LEGOs. Because it understands the "smoothness" of the world, it can easily blend concepts it has never seen together before. The Discrete Brain struggles because it can't "slide" between its checkboxes; if it hasn't seen a specific combination of checkboxes, it gets stuck or produces a "glitchy" mess.

The "Missing Instruction" Problem

The researchers also found that for an AI to be a good "builder," it needs perfect instructions.

Imagine trying to follow a recipe that says, "Add some salt and some spice," but it never tells you exactly how much. You might get close, but you'll probably mess up. The researchers found that if you give the AI "fuzzy" or "incomplete" instructions during training (like saying "red" instead of the exact shade of crimson), the AI loses its ability to combine things creatively later on. It becomes a "copycat" rather than a "creator."

The Fix: Giving the "Sticker" Model a "Slider"

The most exciting part of the paper is that the researchers found a way to "upgrade" the rigid, checkbox-style models.

They gave the Discrete Brain a secondary task: while it was learning to pick the right "stickers," it was also forced to practice "sliding" along a continuous spectrum (a method they called JEPA).

It’s like teaching a child who only uses stickers to also practice painting with watercolors. By doing both, the child learns the precision of the stickers but gains the creative flexibility of the paint. As a result, the model became much better at creating novel, unseen combinations.

Why does this matter?

This isn't just about shapes and colors. This research is a roadmap for building the next generation of AI.

  • In Video/Self-Driving Cars: It helps AI understand that a "left turn at night" is a combination of two known things (turning + darkness), allowing it to handle new weather or lighting conditions safely.
  • In Language: It suggests that AI might think more clearly if it treats "thoughts" as a continuous flow rather than just a sequence of rigid words.

In short: To make AI truly creative and reliable, we need to stop teaching it to memorize the "labels" and start teaching it to understand the "spectrum."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →