New York Smells: A Large Multimodal Dataset for Olfaction
The paper introduces "New York Smells," a large-scale multimodal dataset containing 7,000 image-olfactory pairs from diverse real-world settings, which demonstrates that visual data facilitates cross-modal olfactory representation learning and outperforms traditional hand-crafted features in tasks like retrieval and classification.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where computers can see a picture of a rose and instantly know it smells sweet. For decades, machines have become masters of sight, sound, and even touch, learning to recognize patterns in photos and audio clips with incredible speed. But there is one sense that has remained a mystery to them: smell. While animals like dogs use their noses to navigate the world, find food, and detect danger, machines have struggled to "sniff" anything at all. The main reason isn't that the technology is impossible, but that the data is missing. To teach a computer to smell, you need a massive library of examples where a specific smell is paired with exactly what is causing it. Until now, most smell data was collected in sterile, quiet labs with single objects, which doesn't reflect the messy, complex reality of the real world.
This paper introduces a massive new step forward called "New York Smells." The researchers, a team from Columbia University, Cornell, and Osmo Labs, went out into the wild—literally walking through parks, libraries, gyms, and city streets in New York City. They built a special backpack rig that held a high-definition camera right next to an electronic nose (a device that detects chemicals in the air). As they walked, they recorded thousands of pairs of images and smells, capturing everything from a slice of pizza to a patch of moss, a wooden bench, and a plastic bag. They didn't just take a snapshot; they recorded the smell over time, capturing how the scent of an object changes as the air moves around it. By pairing these real-world smells with the visual data of what they were smelling, they created a dataset with 7,000 unique smell-image pairs, covering 3,500 different objects. This is about 70 times larger than any previous smell dataset.
The team used this new library to teach computers how to understand smells. They tried a method called "contrastive learning," which is like playing a giant matching game. The computer is shown a picture of a book and the smell of a book at the same time, and it learns to link the two. Then, it is tested to see if it can look at a new smell and find the matching picture, or look at a picture and guess the smell. The results were surprising and promising. The computer learned that raw, unprocessed smell data (the messy, real-time signals from the sensor) was much better for learning than the "cleaned-up" summaries scientists usually use. In fact, the computer learned to recognize objects and materials from smell alone, but only because it used the visual signals as a supervisor during the training process. It could even tell the difference between two very similar types of grass growing side-by-side in a park.
The paper suggests that by giving machines a visual "teacher," we can help them learn to smell in the real world, not just in a lab. The researchers found that when the computer used the pictures to guide its learning, it developed a much sharper sense of smell than when it tried to learn from the scent data alone. This doesn't mean the computer can smell like a dog yet, but it proves that linking sight and smell is a powerful way to teach machines about the chemical world. The team is careful to note that their data comes from one city and that their sensors can't detect every chemical, but they have opened a door. They have shown that if we give machines a diverse, real-world library of smells and sights, they can start to build their own understanding of the invisible world around us.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.