CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels
The paper introduces CapFrame, a novel framework that addresses the task of Text-Instructed Viewpoint Grounding (TIVG) in 3D Gaussian Splatting scenes by converting text instructions into geometric pseudo labels through a Retrieve-Translate-Refine pipeline to automatically optimize 6-DoF camera poses for desired frame compositions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine standing inside a room that has been perfectly recreated in a computer, not as a flat picture, but as a three-dimensional space you can walk through. In recent years, scientists have developed ways to build these digital rooms so realistically that they look like photographs and can be viewed instantly from any angle. This technology, known as 3D Gaussian Splatting, has opened the door to virtual reality experiences that feel incredibly lifelike. However, a significant hurdle remains: while the room exists, finding the perfect spot to stand and take a picture of a specific scene is still a manual, tedious process. If a user wants a view where a teddy bear sits on the right side of the frame, facing a sheep on the left, with the camera tilted slightly, they currently have to guess and check, moving the virtual camera back and forth until the composition feels right. Existing tools can find objects, but they struggle to understand the artistic intent behind how those objects should be arranged in a single snapshot.
To solve this problem, a team of researchers has introduced a new method called CapFrame, which allows a computer to automatically find the exact camera position that matches a written description. Instead of just locating an object, the system understands complex instructions about layout, orientation, and camera angle. The researchers tested this approach in 38 different real-world scenes, ranging from indoor rooms to outdoor landscapes, using 135 specific text instructions. They found that their method consistently produced views that aligned much better with the written descriptions than previous techniques, which often failed to get the subject's position or the camera's tilt correct. By turning human language into geometric rules, CapFrame bridges the gap between a simple sentence and a precise, photorealistic image.
The core of this innovation lies in how the computer interprets language. Traditional methods might look for a "teddy bear" and stop there, but CapFrame breaks down a sentence like "a large brown teddy bear sits on the right side, facing a white plush sheep on the left" into a series of specific questions. It asks itself: Is the bear visible? Is it on the right? Is it facing the correct way? To answer these, the system uses a powerful type of artificial intelligence trained to understand both images and text. This AI acts as a judge, scanning through thousands of pre-existing views of the 3D scene to find the ones that come closest to the description. It doesn't just pick the first match; it ranks them based on how well they satisfy every part of the instruction, ensuring the starting point for the next step is already quite good.
Once the system has a promising starting view, it translates the vague words of the instruction into concrete geometric targets. It creates a mental map of where the bear should be in the frame and exactly how the camera should be angled. For instance, if the instruction says the camera should be tilted counterclockwise by 20 degrees, the system converts that into a specific numerical goal for the camera's rotation. It also defines a target area for the bear, ensuring it occupies the right side of the image. These targets act as a guide, telling the computer exactly what the final image should look like before it even begins to move the camera.
The final step is where the magic of mathematics happens, though without any need for the user to understand the equations. The system takes the camera position and begins to nudge it, testing millions of tiny adjustments to see which direction makes the image look more like the description. It does this by calculating the difference between the current image and the target goals it set earlier. If the bear is too far to the left, the system moves the camera right. If the camera is tilted the wrong way, it corrects the angle. This process happens continuously and automatically, refining the view until the image perfectly matches the text instruction. The researchers found that this refinement step was crucial; simply picking a view from the database was not enough, but fine-tuning the position based on the written rules produced a result that felt intentional and precise.
In their experiments, the researchers compared their new method against older techniques that relied on simple guessing or basic search patterns. The results were clear: the older methods often placed the subjects in different spots or failed to capture the correct angle. For example, when asked to show a close-up of a meal with a sandwich behind the main dish, older systems might show the food but miss the specific arrangement. CapFrame, however, successfully arranged the scene exactly as described. The team also asked human volunteers to judge the results, and the participants overwhelmingly preferred the images generated by CapFrame, noting that they felt more natural and better aligned with the instructions.
This work demonstrates that computers can now understand not just what is in a scene, but how a human photographer would choose to frame it. By combining the ability to read complex descriptions with the power to adjust a 3D camera in real-time, CapFrame makes it possible to explore virtual environments using nothing but words. It transforms the experience from a manual search into an intuitive conversation, where a user can simply ask for a view and receive a perfectly composed image. While the system still relies on the quality of the initial 3D scene and the clarity of the text, it represents a significant step forward in making digital worlds easier to navigate and more responsive to human intent.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.