← Latest papers
💻 computer science

HACI-Net: A Hybrid Attention and Channel Interaction Network for Food Image Recognition

This paper proposes HACI-Net, a ConvNeXt-based hybrid network integrating spatial, coordinate, and channel attention mechanisms with channel interaction modules to effectively address challenges in food image recognition, achieving state-of-the-art accuracy on the VireoFood-172 and ChineseFoodNet datasets.

Original authors: Zhiyong Xiao, Chaoliang Liu, Zhaohong Deng

Published 2026-09-17
📖 4 min read☕ Coffee break read

Original authors: Zhiyong Xiao, Chaoliang Liu, Zhaohong Deng

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Every day, billions of people make choices about what to eat, choices that ripple through their health, energy, and long-term well-being. For decades, tracking these choices has relied on human memory and manual logging, a process prone to error and fatigue. In recent years, scientists have turned to computers to solve this problem, teaching machines to look at a photograph of a meal and identify exactly what is on the plate. This field, known as food image recognition, aims to automate dietary monitoring, helping people manage weight, prevent chronic diseases, and understand their nutritional intake without the burden of constant manual entry. However, teaching a computer to distinguish between two very similar dishes is notoriously difficult. A bowl of noodles might look different depending on the lighting, the angle of the camera, or the clutter of a spoon and napkin in the background. Furthermore, the visual clues that humans use to tell foods apart—such as the specific texture of a sauce or the arrangement of ingredients—are often scattered across different layers of visual data, making it hard for standard computer programs to piece them together into a clear picture.

To address these persistent challenges, researchers at Jiangnan University have developed a new system called HACI-Net. This system is designed to look at food images the way a careful observer might: by focusing intently on the most important parts of the dish while ignoring the distracting background, and by carefully combining different visual clues to form a complete understanding. The team built their system upon a modern foundation known as ConvNeXt, which is already quite good at recognizing patterns in images. However, they found that this foundation alone was not enough to handle the subtle differences between similar food categories. To fix this, they added two specific tools to the system. The first is a hybrid attention mechanism, which acts like a spotlight. It scans the image to find the most significant areas, such as the main portion of food, and dims out the noise of the background, like plates or tablecloths. This ensures the computer is looking at the food itself, not the environment surrounding it.

The second tool is a channel interaction module, which helps the system connect different types of visual information. When a computer analyzes an image, it breaks the picture down into many separate channels, each capturing a specific aspect like color, shape, or texture. In older systems, these channels often worked in isolation, failing to share what they learned with one another. The new module forces these channels to talk to each other, allowing the system to realize that a specific red color and a specific smooth texture belong to the same ingredient. By mixing these signals together, the system creates a much richer and more accurate description of the food. The researchers tested this new approach on two large collections of food photographs: one containing 172 different types of dishes from various dining environments, and another containing 208 common Chinese dishes that are often visually very similar to one another.

The results showed that this combined approach works significantly better than previous methods. On the first collection of images, the new system correctly identified the food 94.17% of the time. On the more difficult second collection, where dishes look nearly identical, it achieved an accuracy of 85.46%. These numbers represent an improvement over other leading computer models, which struggled to reach these levels of precision. The researchers verified this success by creating visual heatmaps, which show exactly where the computer was looking when it made its decision. These maps revealed that the new system focused sharply on the food items, whereas older models often got distracted by background objects. While the system is currently quite large and requires powerful computers to run, the researchers acknowledge that future work will focus on making it smaller and faster so it can eventually run on everyday devices like smartphones. For now, the study demonstrates that by teaching computers to pay better attention and to share information more effectively, we can build tools that understand our diets with a level of detail previously thought difficult to achieve.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →