FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
This paper introduces FactorJEPA, a novel architecture that decomposes future predictions into distinct layout, agent, and interaction channels to better model the chaotic, crowded dynamics of Global South urban environments, supported by the release of the large-scale DENSEWORLD dataset and demonstrating superior robustness and accuracy over existing monolithic world models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to understand the world. You don't just want it to recognize a cat or a car; you want it to predict what happens next. If a ball rolls toward a wall, the robot should know it will bounce. If a crowd of people is walking, the robot should guess how they will weave around each other without bumping. This field of science is called "world modeling." It's like giving an AI a crystal ball, but instead of magic, it uses math and video to guess the future.
For a long time, scientists tested these robots in very neat, orderly worlds. Think of a highway with clear white lines, where cars stay in their lanes and everyone follows the rules perfectly. It's like playing a video game on "Easy Mode." But the real world isn't like that. In many bustling cities around the world, traffic is a chaotic dance. There are no strict lanes, people walk where they want, rickshaws squeeze past buses, and animals wander through the streets. It's "Hard Mode." The big question scientists have been asking is: Can our current AI models handle this messy, crowded reality, or do they only work when everything is tidy?
This paper, titled FactorJEPA, tackles that exact problem. The researchers realized that standard AI models were getting confused by the chaos of crowded cities, so they built a brand-new dataset called DENSEWORLD. They collected 1,000 hours of video from 22 different cities, capturing everything from busy markets and narrow alleyways to flyovers and beaches. They found that existing AI models, which usually work great on orderly highways, struggled to predict what would happen next in these crowded scenes. The models would get the general idea but miss the specific details of how people and vehicles interact.
To fix this, the team invented a new method called FactorJEPA. Instead of trying to guess the entire future in one giant, messy lump of data, they taught the AI to break the prediction down into three separate "channels," like a chef separating ingredients before cooking a complex dish:
- Layout: The static stuff that doesn't move much, like buildings, roads, and walls.
- Agents: The moving things, like people, cars, and rickshaws.
- Interactions: How those moving things relate to each other, like a pedestrian stepping aside for a bus or two bikes weaving past one another.
By separating these three things, the AI could learn much better. It's like trying to understand a crowded party. If you try to memorize the whole room at once, you get overwhelmed. But if you focus on the room's shape first, then the people, and finally how they are talking to each other, it becomes much easier to guess who will move next.
The results were impressive. When tested on their new chaotic city videos, this new "separated" AI was much better at predicting the future than the old "all-in-one" models. It made fewer mistakes about where things would be, it was better at understanding what would happen if a car suddenly swerved (a test called "intervention"), and it stayed accurate even when parts of the video were hidden or blurry. The paper shows that by organizing how the AI thinks about the world—breaking it down into layout, agents, and interactions—we can finally build robots that understand the messy, crowded, and wonderful reality of human cities.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.