← Latest papers
🤖 AI

Benchmark for Assessing Olfactory Perception of Large Language Models

This paper introduces the Olfactory Perception (OP) benchmark, a dataset of 1,010 questions designed to evaluate large language models' ability to reason about smell, revealing that current models rely primarily on lexical associations rather than structural molecular reasoning and achieve a maximum accuracy of 64.4% while showing improved performance through multilingual aggregation.

Original authors: Eftychia Makri, Nikolaos Nakis, Laura Sisson, Gigi Minsky, Leandros Tassiulas, Vahid Satarifard, Nicholas A. Christakis

Published 2026-04-02
📖 5 min read🧠 Deep dive

Original authors: Eftychia Makri, Nikolaos Nakis, Laura Sisson, Gigi Minsky, Leandros Tassiulas, Vahid Satarifard, Nicholas A. Christakis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot that has read almost every book, website, and article ever written. It can write poetry, solve complex math problems, and even write computer code. But there's one thing it's never really "experienced": smell.

You can't put a smell into a book. You can't describe the scent of a fresh strawberry or a wet dog with perfect accuracy using just words. This paper introduces a new test, called the Olfactory Perception (OP) Benchmark, to see if these super-smart robots (Large Language Models, or LLMs) can actually "guess" what things smell like, or if they are just guessing based on the words they've seen before.

Here is the breakdown of what they did and what they found, using some simple analogies.

1. The Test: A "Smell Quiz" for Robots

The researchers created a massive quiz with 1,010 questions. Think of it like a final exam for a smell school. The questions cover different levels of difficulty:

  • The Easy Stuff: "Does this molecule smell, or is it odorless?" (Like asking, "Is this rock wet?")
  • The Middle Stuff: "Is this smell more like a rose or a lemon?" or "Is this smell strong or weak?"
  • The Hard Stuff: "Here is a mix of 10 chemicals. Does it smell like a strawberry or a fish?" or "Which specific biological sensors in a human nose would this molecule trigger?"

2. The Twist: The "Name" vs. The "Blueprint"

This is the most clever part of the study. The researchers asked the robots the same questions in two different ways:

  • Method A (The Name): They gave the robot the common name, like "Vanilla."
  • Method B (The Blueprint): They gave the robot the chemical formula (a string of letters and numbers called SMILES), which is like the molecule's DNA or architectural blueprint.

The Analogy: Imagine you are trying to guess what a cake tastes like.

  • Method A is telling you, "This is a Chocolate Cake."
  • Method B is handing you a list of ingredients and chemical structures without saying what the cake is.

3. The Results: Cheating vs. Understanding

The results were surprising and told a clear story:

  • The Robots are "Word Cheaters": When the robots were given the names (e.g., "Vanilla"), they did pretty well. They knew that "Vanilla" usually smells sweet.
  • The Robots are "Blind to Blueprints": When the robots were given the chemical blueprints (the SMILES strings), their performance dropped significantly.

What this means: The robots aren't actually "understanding" how molecules work. They aren't looking at the shape of the molecule and thinking, "Ah, this shape fits into a nose sensor that smells like mint." Instead, they are just memorizing that the word "Mint" is often associated with certain smells. They are relying on lexical associations (word connections) rather than structural reasoning (understanding the physical object).

It's like a student who memorized the answer key for a test but doesn't actually understand the math. If you change the numbers but keep the question the same, they fail.

4. The "Mixing Bowl" Problem

The hardest part of the test was Mixture Similarity.

  • The Task: "Here is a bowl of chemicals A, B, and C. Here is a bowl of chemicals X, Y, and Z. Do they smell the same?"
  • The Reality: In the real world, two completely different bowls of chemicals can smell identical (like two different recipes for the same soup).
  • The Robot Failure: The robots failed miserably here. They tried to count the ingredients. If the bowls didn't have the exact same ingredients, the robots said, "These smell totally different!" They couldn't grasp the concept that the whole can smell different from the sum of its parts.

5. The Language Barrier

The researchers also tested the robots in 21 different languages.

  • The Finding: Smell is tricky in some languages. Some languages have very specific words for smells, while others (like English) are a bit vague.
  • The Good News: When they combined the answers from all 21 languages, the robots got smarter! It's like having a team of experts where one speaks French, one speaks Japanese, and one speaks German. By pooling their knowledge, they could guess the smell better than any single language could.

6. The Safety Glitch

One funny (but serious) side note: Some of the smartest robots refused to answer questions about dangerous chemicals (like nerve agents) because their safety filters kicked in. They wouldn't say, "Yes, this smells bad," because they were programmed to be helpful and harmless. This shows that even in a science test, the robot's "conscience" can get in the way of the answer.

The Big Takeaway

Currently, these AI models are great at remembering words but terrible at understanding the physical world of smell.

  • They know that "Coffee" smells like coffee because they've read the word "coffee" a million times.
  • They don't know why a specific chemical structure smells like coffee.

The Future:
The authors say this is just the beginning. To make AI truly "smell," we need to teach them to look at the molecular blueprints, not just the names. We need them to move from being word-memorizers to molecular-understanders. Until then, if you ask an AI what a new, weird chemical smells like, it's just guessing based on the words it's seen, not the science it understands.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →