← Latest papers
🤖 AI

It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

This paper examines the complexities of integrating AI into augmentative and alternative communication (AAC) systems across six problem spaces, arguing for more robust, intersectional evaluation methods that capture the nuanced desires of users beyond current metrics.

Original authors: Blade Frisch, Will Wade, Dylan Gaines, Michelle Kinsella, Betts Peters, Tamara Broderick, Keith Vertanen

Published 2026-06-24
📖 6 min read🧠 Deep dive

Original authors: Blade Frisch, Will Wade, Dylan Gaines, Michelle Kinsella, Betts Peters, Tamara Broderick, Keith Vertanen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Why "Good Enough" Isn't Good Enough

Imagine you are trying to tune a radio to find your favorite song. In the world of standard computer software, engineers usually just check if the signal is clear and the volume is loud. They use a simple ruler to measure "speed" and "accuracy."

But for people who use AAC (Augmentative and Alternative Communication)—devices that help them speak when they can't use their natural voice—this simple ruler doesn't work. These users are not just "radio signals"; they are complex, intersectional human beings with unique identities, changing moods, and specific social needs.

The authors argue that if we only measure AI-powered AAC devices by how fast they type or how few typos they make, we are like a chef who only tastes the salt in a soup and ignores the flavor, texture, and whether the person eating it actually enjoys the meal. We risk building systems that are technically perfect but feel wrong to the user.

The Core Problem: The "One-Size-Fits-All" Trap

The paper suggests that current AI often tries to fit everyone into a single, "standard" box. This is like trying to force a square peg into a round hole.

  • The Issue: AI models often optimize for a "hypothetical average user."
  • The Result: This ignores the fact that a person's identity is a mix of many things (race, gender, disability, culture). When AI ignores this mix, it can accidentally make the user feel misunderstood or "othered."
  • The Analogy: Imagine a translator who only knows how to speak in a stiff, formal business tone. If you try to tell a funny joke to a friend or whisper a secret to a doctor, the translator ruins the moment because it doesn't understand the context.

Six Areas Where AI Can Help (and Where We Need to Be Careful)

The authors explore six specific areas where AI could make AAC better, but they warn that we need new ways to measure success in each one.

1. Speed vs. Accuracy (The Race Car vs. The GPS)

  • The Challenge: Speaking is fast (like a race car), but typing on an AAC device is slow. AI can predict words to speed things up, like a GPS suggesting a shortcut.
  • The Risk: If the GPS suggests a shortcut that leads to a dead end, you waste time. If the AI predicts the wrong word, the user has to stop and fix it.
  • New Way to Measure: Don't just count how many words per minute were typed. Ask: "Did the user feel like they were saying what they meant to say?" Sometimes, a slightly slower prediction that captures the feeling of the sentence is better than a fast, literal one.

2. Physical and Mental Effort (The Backpack)

  • The Challenge: Using these devices can be exhausting. Some users have to tap a switch many times just to select one letter. It's like carrying a heavy backpack up a mountain.
  • The Risk: If AI tries to help by showing too many options, the user might get mentally tired trying to choose. If it shows too few, they might have to tap too many times physically.
  • New Way to Measure: We need to measure the "weight" of the backpack. Does the AI actually reduce the number of taps needed? Does it reduce the mental stress of searching? We need to ask the user if they feel tired, not just count their button presses.

3. Sounding Like Yourself (The Voice Mask)

  • The Challenge: AAC voices often sound robotic and flat, like a text-to-speech robot. But humans use tone, accent, and emotion to show who they are.
  • The Risk: If the AI makes a user sound like a generic robot, they lose their personality. It's like wearing a mask that hides your face.
  • New Way to Measure: Can the AI help the user sound like them? Does the voice sound like an extension of their soul? We need to ask friends and family: "Does this sound like the person you know?"

4. Switching Codes and Contexts (The Chameleon)

  • The Challenge: We all talk differently to our boss than we do to our best friend. We might switch languages or slang depending on the room we are in.
  • The Risk: Most AAC systems are stuck in one mode. They can't switch from "Job Interview Mode" to "Coffee Shop Mode."
  • New Way to Measure: Does the AI know when to be formal and when to be casual? Can it help the user fit in naturally with different groups of people?

5. Joining the Conversation (The Dance Floor)

  • The Challenge: Conversation is a dance. You have to know when to step in, when to nod, and when to say "uh-huh." Because AAC is slow, users often miss their turn to dance.
  • The Risk: If the AI helps them jump in too early or too late, it feels awkward. It's like a dance partner who steps on your toes.
  • New Way to Measure: Did the user get to participate in the flow of the conversation? Did their communication partners feel like they were talking with the user, or just waiting for the machine to finish?

6. Changing Needs Over Time (The Growing Pains)

  • The Challenge: People's needs change. A user might be energetic in the morning but tired in the afternoon. Or, as a condition progresses, they might need a different way to control the device.
  • The Risk: A static system is like a shoe that never stretches. Eventually, it will pinch and hurt.
  • New Way to Measure: Can the system adapt? If the user gets tired, does the AI automatically offer bigger, easier predictions? We need to watch how the system behaves over weeks and months, not just in a one-hour lab test.

The Solution: A New Toolkit for Evaluation

The paper concludes that we cannot rely on a single "score" to judge these systems. Instead, we need a pluralistic toolkit—a mix of different tools to get the full picture.

  • The Old Way: Just count the errors and speed.
  • The New Way: Combine the numbers (technical metrics) with the stories (human interviews, diaries, and user feedback).

The Final Analogy:
Evaluating an AI-powered AAC system is like judging a new car.

  • The Technical Metric: Checking the engine horsepower and fuel efficiency.
  • The Human Metric: Asking the driver, "Is the seat comfortable? Does the radio feel intuitive? Do you feel safe and in control?"

The authors say we must stop treating AAC users as data points to be optimized and start treating them as partners. We need to design and test these systems with the users, ensuring the technology serves their unique, complex, and beautiful human identities, rather than forcing them to fit into a machine's narrow definition of "correct."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →