GazeOnce360: Fisheye-Based 360{\deg} Multi-Person Gaze Estimation with Global-Local Feature Fusion
This paper introduces GazeOnce360, a novel end-to-end model for estimating 3D gaze directions of multiple people from a single upward-facing fisheye camera, supported by the new MPSGaze360 synthetic dataset and a dual-resolution architecture that fuses global context with local eye features to handle severe fisheye distortion.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are sitting at a round table in a busy coffee shop. You want to know exactly what everyone around you is looking at—maybe to see who is interested in the menu, who is distracted, or who is making eye contact with you.
In the past, to do this, you would need a team of people with cameras standing around the table, or a complex system of mirrors and lenses. It would be expensive, clunky, and prone to mistakes.
Enter "GazeOnce360."
Think of this new technology as a super-powered, all-seeing eye placed right in the middle of your table, looking straight up at the ceiling. It's a special "fisheye" camera (like the ones used for security or 360-degree photos) that can see everyone in a full circle around it.
Here is how the paper explains this breakthrough, broken down into simple concepts:
1. The Problem: The "Funhouse Mirror" Effect
The main challenge is that fisheye cameras are like funhouse mirrors. They stretch and warp everything.
- If you look at a person standing far to the side in a normal photo, their face looks normal.
- In a fisheye photo, that same face looks stretched, squished, and rotated.
- Previous methods tried to fix this by taking the warped image, cutting it into pieces, straightening them out, and then guessing where people are looking. This is like trying to solve a puzzle by taking it apart, fixing each piece, and then putting it back together. It's slow, and if you make a mistake on one piece, the whole picture is wrong.
2. The Solution: The "One-Step" Magic Trick
The authors built a new AI model called GazeOnce360. Instead of the slow, multi-step puzzle method, this model is like a magician who sees the whole picture at once.
- It looks at the warped, stretched fisheye image directly.
- It doesn't try to "fix" the image first; instead, it learns to understand the distortion as a natural part of the view.
- It instantly tells you: "That person on the left is looking at the coffee cup," and "That person on the right is looking at you."
3. How It Works: The "Dual-Lens" Strategy
To be both fast and accurate, the model uses a clever dual-resolution trick. Imagine a security guard who has two ways of looking at a crowd:
- The Wide-Angle View (Low-Res): The guard looks at the whole room from a distance. They can see where everyone is standing and how the group is arranged, but they can't see the details of their eyes.
- The Zoom-In View (High-Res): The guard then zooms in on each person's face to see exactly where their pupils are pointing.
GazeOnce360 combines these two views. It uses the "Wide-Angle" to find the people and the "Zoom-In" to read their eyes. It then fuses this information together so it knows exactly who is looking at what, without getting confused by the warped edges of the fisheye lens.
4. The Secret Sauce: Rotational Convolution
Standard AI is used to seeing things upright. If you rotate a picture of a cat, a normal AI might get confused. But in a 360-degree fisheye view, people are rotated in every direction.
The authors gave their AI a special tool called Rotational Convolution.
- Analogy: Imagine a standard flashlight that only shines straight ahead. If you turn your head, the light doesn't follow.
- The New Tool: This is like a 360-degree spotlight that rotates with the object. No matter how the face is twisted or stretched by the fisheye lens, the AI's "spotlight" rotates to match it, ensuring it always sees the eyes clearly.
5. The Training Ground: The "Virtual Reality" Classroom
You can't easily teach a robot to do this with real people because it's hard to know exactly where a human is looking (you'd have to ask them, and they might lie or forget).
So, the authors built a massive virtual world (using a game engine called Unreal Engine).
- They created thousands of digital people with different skin tones, ages, and clothes.
- They placed them around a virtual table and programmed them to look in specific directions.
- Because it's a computer simulation, the computer knows exactly where every pupil is looking, down to the pixel.
- They trained the AI on this "fake" data, and surprisingly, it learned so well that it works perfectly on real-world photos too!
Why Does This Matter?
This isn't just a cool tech demo; it has real-world uses:
- Smart Offices: A robot on a table could know if the whole team is bored or engaged during a meeting.
- Service Counters: A bank or hotel desk could see if a customer is looking at a brochure or waiting for help.
- Virtual Reality: It could make VR interactions feel more natural by understanding where you and others are looking in a 360-degree space.
In short: GazeOnce360 is a smart, fast, and flexible system that turns a distorted, 360-degree view into a clear understanding of human attention, all without needing a team of cameras or a complicated setup. It's like giving a computer the ability to read the room, no matter how the room is arranged.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.