TouchAI: Exploring human-AI perceptual alignment in touch through language model representations
This paper investigates the perceptual alignment between large language models and human touch experiences through a "Guess What Textile" task, revealing that while some degree of alignment exists, it varies significantly across different textile types and often fails to match human subjective perceptions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Can AI "Feel" What We Feel?
Imagine you are trying to teach a robot what a silk scarf feels like versus a pair of jeans. You can't give the robot hands, so you have to describe it using words. You say, "The silk is smooth and slippery, like water," and "The jeans are rough and stiff, like sandpaper."
This study asked a simple question: If a human describes how something feels, can a smart computer (an AI) understand that description well enough to guess exactly what the object is?
The researchers called this project "TouchAI." They wanted to see if the AI's "mental map" of how things feel matches up with our own human map.
The Experiment: The "Blind Guessing Game"
To test this, the researchers set up a game called "Guess What Textile?"
- The Setup: A participant sat at a desk with a black box in front of them. They couldn't see inside, only reach in.
- The Task: Inside the box were two pieces of fabric. One was a "reference" (the AI knew what this was, like a piece of cotton). The other was a "target" (the AI didn't know this one).
- The Interaction: The participant touched both fabrics. They then had to describe the difference between them to the AI.
- Example: "The target feels softer and smoother than the cotton reference."
- The AI's Turn: The AI listened to the description, processed the words, and made a guess: "I think the target is Silk Satin."
- The Result:
- If the AI was right, the game ended.
- If the AI was wrong, the "Silk Satin" became the new reference, and the participant had to describe the target again, comparing it to the silk. The AI tried again. They could do this up to five times.
What Did They Find?
The results were a mix of "not bad" and "very confused."
1. The AI is mostly guessing in the dark.
Overall, the AI only got the right answer 22.5% of the time. That is barely better than if the AI had just closed its eyes and picked a random fabric from a hat. This suggests that current AI models don't really "know" what touch feels like; they are mostly just guessing based on how often they've seen those words in books.
2. The AI has a "Favorite" fabric.
Here is where it gets interesting. The AI wasn't equally bad at everything.
- The Winner: When the target was Silk Satin, the AI got it right 100% of the time.
- The Losers: When the target was Cotton Denim (jeans), the AI got it right 0% of the time.
Analogy: Imagine a student taking a test. They got every question about "Gold" right because they memorized a poem about gold. But they got every question about "Sand" wrong because they never read about sand. The AI knows the "Gold" (Silk) because it appears often in stories and movies as "soft and shiny." It doesn't know the "Sand" (Denim) because people rarely write detailed stories about how rough jeans feel.
3. The "Mental Map" is skewed.
The researchers looked at how the AI organized these fabrics in its "brain" (its data space).
- Human Map: We group fabrics by how they feel. Silk and a soft synthetic might be neighbors because they both feel smooth.
- AI Map: The AI groups fabrics by what they are called or what category they belong to. It might put all "cottons" together and all "silks" together, even if a specific cotton feels very different from another.
The Metaphor: Think of the AI's brain like a library where books are sorted by the color of the cover (the name of the fabric) rather than the genre (how it feels). Humans sort by genre (feeling); the AI sorts by cover color (name).
Why Does This Matter?
The paper explains that this happens because touch is hard to write about.
- Vision is easy: We have a universal language for colors. "Red" is always #FF0000. We can describe a red apple very precisely.
- Touch is messy: Words like "soft," "rough," or "buttery" are vague. One person's "soft" is another person's "medium." Also, people rarely write long descriptions of how their jeans feel in their daily emails or novels.
Because the AI was trained on text (books, websites, articles), it learned a lot about how things look and sound, but very little about how things feel. It's like trying to learn what an elephant feels like by only reading a book about elephants, but never actually touching one.
The Bottom Line
The study concludes that while AI is getting very good at understanding words, it is still bad at understanding feelings.
- It can guess "Silk" because the word "Silk" is strongly linked to "soft" in its training data.
- It fails at "Denim" because the connection between the word "Denim" and the feeling of "roughness" isn't strong enough in the text it read.
The researchers say that for AI to truly help us in the future (like a robot that picks out clothes for you), it needs to learn more than just words. It needs to understand the physical world, not just the description of it. Until then, if you ask an AI to pick a fabric that feels "just right," it might just guess based on the most popular word it knows, not what actually feels good.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.