Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
This paper introduces Gazette, a novel framework that leverages multimodal large language models and synthetic "think-aloud" transcripts to decode human gaze scanpaths into free-form natural language descriptions of goals, moving beyond traditional categorical decoding to capture the open-ended nuances of human intentions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine your eyes are like a flashlight beam in a dark room. When you look at something, that beam doesn't just sit still; it darts around, landing on different spots for tiny fractions of a second. This pattern of movement is called a "scanpath." For a long time, scientists have tried to figure out what you are thinking just by watching where your flashlight beam lands. Usually, they've treated this like a multiple-choice quiz. They'd ask, "Is the person looking for a cat or a dog?" and the computer would pick one answer from a fixed list. But human thoughts are messy and creative. Sometimes you aren't just looking for a "dog"; you are looking for "the scruffy golden retriever wearing a red bandana that is hiding behind the couch." The old way of asking questions was too simple to capture that kind of detail. This is where a new field called "gaze decoding" comes in, trying to translate those eye movements into the rich, open-ended stories happening inside our heads.
Enter a new research paper that introduces a clever new way to solve this puzzle. The authors, working at Stony Brook University and Adelaide University, have built a system called Gazette. Instead of forcing the computer to pick from a multiple-choice list, Gazette is designed to write a free-form story about what a person is looking for, using natural language just like a human would. Think of it as upgrading from a rigid checklist to a creative writer who can describe exactly what you see.
The big problem the researchers faced was that everyone's eyes move a little differently, even when they are looking for the same thing. If ten people are asked to find a "red car," ten different people might look at the car in ten slightly different ways. Some might look at the wheels first, others at the windows. If you just show a computer one person's eye movements, it gets confused by all those personal habits. It's like trying to guess a song by listening to one person humming it off-key; you might hear the tune, but the noise makes it hard to be sure.
To fix this, the team came up with a smart trick involving a "think-aloud" strategy. They realized that while everyone's eye movements are unique, the reason they are moving is the same. So, they used a powerful AI (GPT-4) to look at the eye movements of many different people all trying to find the same object. The AI acted like a detective, ignoring the individual quirks and finding the common pattern—the "secret recipe" for how humans look for that specific thing. The AI then wrote a "think-aloud transcript," which is basically a story explaining the strategy: "First, look at the big shapes, then zoom in on the red parts."
They taught their new model, Gazette, to read these AI-generated stories along with the eye movements. This helped the model learn to separate the "goal" (what the person is looking for) from the "noise" (the person's unique way of looking). When they tested it, Gazette was much better at guessing the goal than previous methods. For example, when asked to find a specific car in a picture, Gazette didn't just say "car." It could say, "The black car in the top right corner," and it did this with high accuracy, even when the task was tricky.
The paper shows that by teaching the computer to understand the strategy behind the eye movements, rather than just the movements themselves, we can decode human intentions with much more detail. While the system works best when people agree on what they are looking for (like finding a specific object), it suggests that we are getting closer to a future where computers can understand our goals just by watching where we look, opening doors for better assistive technologies and a deeper understanding of how we interact with the world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.