Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video
SAGE is a scalable framework that adapts 3D geometric foundation models using unannotated internet videos by employing a hierarchical mining pipeline that combines sparse SfM-based structural guidance with dense differentiable 3D Gaussian rendering for hybrid supervision.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a child how to draw a perfect 3D map of a room.
To do this perfectly, you would normally need to give them a professional architect’s blueprint (this is what scientists call "labeled 3D data"). But there’s a problem: blueprints are expensive, hard to find, and only exist for a few thousand rooms. If you only show the child a few blueprints, they’ll be great at drawing those specific rooms, but if you take them into a new park or a different house, they’ll be totally lost.
This paper introduces SAGE, a new way to teach these "digital artists" (3D Foundation Models) using something we have in infinite supply: YouTube videos.
Here is how SAGE works, explained through three simple metaphors:
1. The Problem: The "Blindfolded Architect"
Current 3D AI models are like architects working in the dark. They have seen a few high-quality blueprints, but they haven't seen the "real world." When they try to reconstruct a scene from a random video, they get confused because videos don't come with blueprints—they are just flat, moving pictures. If you try to teach them using just "guesses" of what the depth is, the AI gets "dizzy" and starts making messy, distorted shapes.
2. The Solution: The Three-Layer Teacher (SAGE)
Instead of giving the AI a blueprint, SAGE acts like a multi-layered teacher that uses the video itself to provide "clues."
- The Skeleton (Sparse Geometric Anchors): Imagine watching a video of a person walking through a forest. Even if you don't have a map, you can see the big, solid things: a massive tree trunk or a large rock. SAGE uses a tool (called SfM) to find these "big, solid landmarks" first. It tells the AI, "I don't know exactly where every leaf is, but I know for sure that this giant rock is right here." This prevents the AI from getting lost in the details.
- The Skin (Dense Differentiable Consistency): Once the "skeleton" is set, the AI needs to fill in the gaps (the grass, the walls, the furniture). SAGE uses a technique called "Gaussian Splatting," which acts like a digital spray paint. It tells the AI, "If you move your camera slightly to the left, the wall should still look like a wall." By constantly checking if the "painted" 3D scene looks consistent from different angles, the AI learns to fill in the fine details smoothly.
- The Memory Guard (Regularization): When you teach someone a new skill (like dancing) using a new method (like a video), they might forget how to do their old skill (like walking). This is called "catastrophic forgetting." SAGE keeps a tiny "cheat sheet" of the original professional blueprints to remind the AI, "Hey, don't forget the basic rules of geometry while you're learning from these wild YouTube videos!"
3. The Result: The "Super-Learner"
The researchers tested SAGE by feeding it 10,000 video clips—ten times more than what is usually used.
The result? The AI became a "super-learner." Because it practiced on so much diverse footage, it didn't just get better at the videos it saw; it became much better at reconstructing entirely new places it had never encountered before. It reduced errors by up to 42% compared to the old way.
Summary in a Nutshell
Before SAGE: AI was a student studying only from a few expensive textbooks. It was smart but narrow-minded.
With SAGE: AI is a student who has watched millions of hours of travel vlogs. It uses the "big landmarks" and "visual consistency" of those videos to build a deep, intuitive understanding of the 3D world, making it a master of reconstruction in almost any environment.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.