← Latest papers
💬 NLP

Your Mouse and Eyes Secretly Leak Your Preference: LLM Alignment using Implicit Feedback from Users

This paper introduces the IFLLM dataset, which leverages implicit user feedback from mouse trajectories and eye gaze to train reward models that significantly improve LLM alignment accuracy and response quality compared to traditional text-based methods.

Original authors: Haw-Shiuan Chang, Jeffrey Gomez, Mehul Patwari, Aryan Sajith, Hamed Zamani

Published 2026-06-19
📖 4 min read☕ Coffee break read

Original authors: Haw-Shiuan Chang, Jeffrey Gomez, Mehul Patwari, Aryan Sajith, Hamed Zamani

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to write good stories. Usually, you have to stop the robot, read what it wrote, and explicitly tell it, "This is good," or "This is bad." This is like asking a student to raise their hand and say, "I understand!" after every sentence. It's slow, expensive, and most people get tired of doing it.

This paper proposes a different, sneakier way to teach the robot: watch what the student does, not just what they say.

Here is the breakdown of the research using simple analogies:

1. The Problem: The "Silent Majority"

Most people don't bother to click "thumbs up" or "thumbs down" on AI answers. They just read and move on. The researchers realized that while people are silent with their words, they are very loud with their actions. They are constantly moving their mouse and looking at the screen. These actions are "implicit feedback"—clues that reveal what the user actually likes without them having to stop and type a review.

2. The Experiment: The "Eye-and-Mouse" Lab

To prove this works, the team built a special website (like a digital playground) where 59 workers from Amazon Mechanical Turk chatted with AI models.

  • The Setup: As the workers asked questions and read answers, the website secretly recorded two things:
    1. Where their eyes looked (using a standard webcam, not expensive lab equipment).
    2. Where their mouse moved.
  • The Dataset (IFLLM): They collected over 1,300 conversations. This is the first time anyone has gathered both eye movements and mouse trails specifically for judging AI answers in a real conversation.

3. The Discovery: The Mouse is the "Scrolling Signal"

The researchers analyzed the data and found some fascinating patterns:

  • The "Skim" vs. The "Deep Dive": When an answer was short, people's eyes and mouse didn't move much together. But when the answer was long, the mouse became a perfect mirror of the eyes. Why? Because to read a long text, you have to scroll. If a user scrolls all the way to the bottom, it means they are interested. If they stop halfway, they might be bored.
  • The "Mouse" is the MVP: Surprisingly, the mouse movement was often a better clue than the eye movement, especially for long answers. The mouse acts like a "commitment meter." If you are willing to drag your mouse to scroll through a long story, you probably like it.

4. The Solution: The "Smart Teacher" (Reward Model)

The team built a new "teacher" (a reward model) for the AI.

  • Old Teacher: Only read the text of the answer and guessed if it was good. It got this right about 55% of the time (basically a coin flip).
  • New Teacher: Looked at the text plus the mouse and eye data. It got this right about 64% of the time.
  • The Result: When they used this "New Teacher" to train the AI (a process called DPO), the AI's answers improved significantly more than when using the old teacher. In fact, the improvement was nearly three times better than before.

5. The Big Picture: A Self-Improving Loop

The paper argues that we don't need to ask users to stop and rate every answer. We can just watch how they interact.

  • The Analogy: Imagine a librarian who doesn't ask you if you liked a book. Instead, she watches to see if you walk all the way to the end of the book, or if you put it back on the shelf after one page. She uses that behavior to recommend better books next time.
  • The Outcome: By using these "silent signals" (mouse and eyes), we can train AI to be more helpful, accurate, and aligned with what humans actually want, without annoying them with constant surveys.

What the Paper Does Not Claim

  • It does not claim this works for medical diagnosis or clinical therapy.
  • It does not claim that AI will soon read your mind or emotions directly.
  • It does not say this is perfect for everyone; it notes that individual reading habits vary wildly (some people read fast, some slow, some skip around).

In short: The paper shows that your mouse and eyes are secretly voting for what you like. By listening to those silent votes, we can teach AI to be much better at talking to us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →