GazePrior: Zero-Shot AR/VR Eye Tracking via Learned 3D Gaze Reconstruction
The paper introduces GazePrior, a zero-shot eye tracking method that leverages a learned 3D prior to synthesize realistic, diverse, and accurately annotated training data from existing devices, enabling high-performance eye tracking on new AR/VR hardware without the need for costly new data collection.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to read your eyes. To do this accurately, the robot needs to see millions of pictures of human eyes looking in every possible direction, under every possible light, from every possible angle.
Usually, to get these pictures, companies have to build a new camera setup, find thousands of volunteers, and spend months taking photos. This is expensive, slow, and a logistical nightmare every time they design a new pair of smart glasses.
The Problem: The "New Camera" Dilemma
Think of eye-tracking cameras like different pairs of glasses. If you change the glasses (the camera), the way the eye looks changes slightly. A model trained on "Glasses A" often fails when you put it on "Glasses B." Traditionally, to fix this, you'd have to go out and take a whole new set of photos with "Glasses B."
The Solution: GazePrior (The "Universal Eye Blueprint")
The researchers behind GazePrior came up with a clever shortcut. Instead of taking new photos, they built a "Universal Eye Blueprint."
Here is how it works, using a simple analogy:
- The Master Sculptor (The 3D Prior): Imagine a master sculptor who has studied thousands of real people's eyes. This sculptor doesn't just memorize one face; they learn the rules of how eyes, eyelids, and eyebrows move together. They understand that when an eye looks left, the eyelid doesn't just stay still—it shifts slightly. This "sculptor" is a computer model called a 3D Prior. It learns the hidden patterns of how eyes look from every angle, for every person, in every light.
- The Sparse Clues (Old Data): The researchers took existing photos of eyes taken with old camera setups (which are already labeled with the correct "gaze direction"). They didn't need new photos; they just needed these old, sparse clues.
- The Magic Translation (Retargeting): Now, imagine you want to see what those same eyes would look like through the lens of a brand new camera that hasn't even been built yet.
- The "sculptor" takes the old photos.
- It uses its "Universal Eye Blueprint" to reconstruct a perfect 3D model of the eye in that moment.
- It then "renders" (draws) a brand new image of that eye as if it were being seen through the new camera's lens.
Why This is a Big Deal
Previous methods tried to do this by taking a 2D photo and just "warping" or stretching the eyeball to look in a new direction. It's like taking a flat sticker of an eye and trying to twist it; it often looks fake, and the eyelids don't move naturally.
GazePrior is different because it builds a true 3D model. It understands that eyes are round objects and that eyelids are flexible skin. Because of this, when it generates new images for a new camera, the eyelids blink naturally, the eyelashes catch the light correctly, and the reflection on the eye looks real.
The Result
The researchers tested this by training an eye-tracking AI using only these computer-generated images.
- The Test: They tried to make the AI work on a new device (the Meta Aria headset) without ever showing it a single real photo taken by that specific device.
- The Outcome: The AI performed almost as well as if it had been trained on real photos taken by that device. It beat all previous "zero-shot" methods (methods that try to work without new data) by a significant margin.
In a Nutshell
GazePrior is like having a master chef who can recreate a complex dish perfectly using only a few ingredients and a recipe, without needing to go to the market to buy fresh produce every time. It allows companies to design and test new eye-tracking glasses using "virtual" data generated from old data, saving them the cost and time of massive real-world photo shoots.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.