← Latest papers
💬 NLP

Proactive Conversational Assistant for a Procedural Manual Task based on Audio and IMU

This paper presents a privacy-preserving, edge-deployed conversational assistant that utilizes only audio and IMU data to provide real-time, proactive guidance for procedural manual tasks, demonstrating that fine-tuning a language model significantly improves its precision in limiting unnecessary dialogue and its recall in answering user questions.

Original authors: Rehana Mahfuz, Yinyi Guo, Erik Visser, Phanidhar Chinchili

Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Rehana Mahfuz, Yinyi Guo, Erik Visser, Phanidhar Chinchili

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to assemble a complicated piece of furniture or cook a new recipe. Usually, you might try to watch a video tutorial on your phone, but that requires you to stop what you're doing, look at a screen, and it can feel like you're being watched by a camera.

This paper introduces a new kind of digital helper that works differently. Think of it as a "silent coach" that lives on your smartwatch. Instead of watching you with a camera, it listens to the sounds you make (like the whir of a drill or the sizzle of a pan) and feels the movements of your wrist (like the rhythm of screwing or chopping).

Here is how the paper breaks down this invention:

1. The "Ears and Feel" Instead of "Eyes"

Most current helpers use cameras to see what you are doing. The authors say cameras are like a heavy backpack: they use too much battery, are slow, and make people feel like their privacy is being invaded.

Instead, this new assistant is like a detective with super-hearing and a sense of touch. It only uses two things:

  • Audio: The sounds of your task.
  • IMU (Inertial Measurement Unit): The motion sensors in your watch that track how your wrist moves.

This makes the system lightweight, fast, and private because no video of your home is ever recorded or sent to the cloud.

2. The "Traffic Cop" and the "Coach"

The system has two main parts working together, like a team:

  • The Traffic Cop (The Orchestrator): This part counts steps. It knows, "You just finished screwing in three bolts, so the next step is the fourth one." It keeps track of time and counts.
  • The Coach (The Language Model): This is the brain that talks to you. It takes the Traffic Cop's data and turns it into natural conversation.
    • If you are on track: It says, "Great job, now grab the next leg."
    • If you make a mistake: It says, "Wait, you drilled that hole too early! Let's unscrew it and try again."
    • If you ask a question: It answers, "How much longer?" or "Why do I need to stir?"

3. Teaching the Coach to be Quiet and Smart

The researchers found that if you just take a standard, pre-made AI (like a generic chatbot) and ask it to help, it tends to be too chatty. It might say, "Let's proceed!" or "I'm here to help!" even when you don't need to hear anything. It's like a coach who won't stop talking during a timeout.

To fix this, they fine-tuned the AI. Think of this as giving the coach a strict rulebook. They taught it:

  • Be concise: Only speak when necessary.
  • Be accurate: Answer questions correctly.
  • Be proactive: Don't wait for you to ask; tell you what to do next before you get stuck.

The Results of Training:

  • It became 50% better at knowing when not to talk (stopping the unnecessary chatter).
  • It became 150% better at actually answering your questions correctly.
  • It became 55% better at sounding helpful and human according to human judges.

4. Real-World Tests: Furniture and Soup

They tested this on two very different tasks:

  1. Building a Table: This requires counting (how many screws?) and order (screw before drilling).
  2. Making Soup: This requires timing (stir every two minutes) and sequence (add oil before veggies).

The system successfully guided users through these tasks, corrected mistakes (like drilling before screwing), and answered questions, all while running entirely on a local device without needing an internet connection.

5. What It Can't Do (The Limitations)

The paper is honest about what this "ears and feel" system cannot do yet:

  • It can't see ingredients: If you chop a carrot instead of a radish, the sound might be similar, and the watch can't tell the difference. A camera could see that, but this system can't.
  • It assumes you are right-handed: It expects the watch to be on your dominant (usually right) wrist.
  • It's not a bodyguard: If you drop a heavy table on your foot or get burned by hot oil, the system can warn you, but it can't physically stop the accident.
  • It needs specific training: The AI is trained for specific tasks. If you try to use the "furniture" AI to cook soup, it won't work well. You need a model trained for that specific job.

Summary

In short, this paper presents a privacy-friendly, battery-efficient assistant that helps you do manual tasks by listening and feeling your movements rather than watching you. By training the AI to be less chatty and more accurate, they created a tool that feels like a helpful, knowledgeable partner right on your wrist, capable of running on your own device without needing the cloud.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →