SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering
SeGPruner is a novel framework for 3D question answering that significantly enhances inference efficiency by reducing visual tokens by 91% through a dual mechanism of attention-based semantic salience preservation and geometry-guided spatial diversification, thereby maintaining robust reasoning performance while addressing token redundancy in multi-view pipelines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a mystery in a huge, cluttered house. You have a very smart detective (the AI) who needs to answer questions like, "What color is the chair next to the white table?" or "Where is the door?"
To help the detective, you take photos of the house from every possible angle—front, back, left, right, up, and down. You hand the detective all these photos.
The Problem: Too Much Clutter
The problem is that taking photos from every angle creates a massive amount of redundancy.
- If you take a photo of a red chair from the left, and then another from the right, the detective sees the red chair twice.
- If you take a photo of a blank white wall from ten different angles, the detective sees that wall ten times.
The detective's brain (the AI model) gets overwhelmed. It has to process thousands of tiny pieces of information (called "tokens") just to find the one chair or the one door. This makes the detective slow, expensive to run, and sometimes confused by the noise.
The Old Solution: The "Random Shredder"
Previous methods tried to fix this by just shredding (removing) some of the photos randomly or based on simple rules.
- The Flaw: They might throw away the photo of the chair because it looked "boring" in one shot, or they might keep ten photos of the same empty wall because they looked slightly different. The detective ends up with a pile of junk and misses the crucial clues.
The New Solution: SeGPruner (The Smart Editor)
The authors of this paper created SeGPruner, a smart editor that acts like a highly skilled film director. Instead of just cutting random scenes, it uses two specific rules to decide what to keep:
1. The "Star Actor" Rule (Semantic Salience)
First, the editor asks: "What is the main character in this scene?"
- It looks at the photos and identifies the important objects (the chair, the table, the door).
- It says, "No matter what, we must keep the photos of these actors." It ensures the detective never loses the critical evidence needed to answer the question.
- Analogy: It's like making sure you keep the close-up shots of the suspect, even if you have to cut out the shots of the empty hallway.
2. The "Wide-Angle" Rule (Geometric Diversity)
Next, the editor asks: "Do we have a good view of the whole house?"
- Even if we keep the chair, we don't need 50 photos of the same side of the chair. We need to see the chair from the front, the side, and maybe the back.
- This tool uses 3D geometry (like a mental map of the room) to pick photos that show different parts of the room. It ensures the detective gets a 360-degree understanding without the clutter.
- Analogy: It's like a photographer who knows, "I have a great shot of the kitchen from the left, so I'll skip the left shot of the living room and take a shot from the right instead to balance the album."
The Result: A Super-Efficient Detective
By combining these two rules, SeGPruner creates a "highlight reel" of the house.
- It cuts out 91% of the unnecessary photos.
- It makes the detective 86% faster.
- It doesn't lose accuracy. In fact, because the detective isn't distracted by the clutter, it often answers questions better than before.
Why This Matters
Think of it like packing for a trip.
- Old Way: You pack your entire closet, hoping you'll find the right shirt. You end up with a heavy suitcase and can't move.
- SeGPruner Way: You pack only the specific outfits you need for the weather (Semantic) and ensure you have options for different activities (Geometric). You have a light suitcase, you move faster, and you are perfectly dressed for the occasion.
In short, SeGPruner teaches AI how to ignore the noise and focus on the signal, using a 3D map to ensure it sees the whole picture without getting overwhelmed by the details.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.