← Latest papers
💻 computer science

SQCAFusion: Multi-Scale Scene-Query Cross-Attention for Illumination-Adaptive Infrared and Visible Image Fusion

This paper proposes SQCAFusion, an illumination-adaptive infrared and visible image fusion framework that employs multi-scale scene-query cross-attention with learnable scene tokens to dynamically modulate features during early encoding, thereby resolving scene-blindness issues and enhancing fusion quality under varying lighting conditions.

Original authors: Yunfei Chen, Juan Zhang, Yongbin Gao, Bo Huang

Published 2026-09-22
📖 5 min read🧠 Deep dive

Original authors: Yunfei Chen, Juan Zhang, Yongbin Gao, Bo Huang

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of machine vision, cameras are often limited by the physics of light. A standard camera sees the world much like a human eye, capturing rich colors and fine textures, but it struggles when the sun goes down or smoke fills the air. In contrast, thermal cameras detect heat rather than light, allowing them to see living things clearly even in total darkness, yet they often produce blurry, gray images that lack the sharp details of a normal photograph. For decades, engineers have tried to combine these two types of images into a single, perfect picture that holds the clarity of a daytime photo and the heat-sensing power of a night-vision device. This process, known as image fusion, is vital for everything from autonomous vehicles navigating foggy streets to military surveillance in hostile environments. However, a persistent problem has plagued these attempts: most computer programs designed to merge these images are "scene-blind." They treat a bright, sunny day and a pitch-black night exactly the same way, failing to realize that the rules for combining the images should change depending on the lighting. This blindness often leads to results where the darkness of the night sky amplifies static and noise, or where the bright colors of a sunny day get washed out, leaving the final image muddy and confusing.

A team of researchers at the Shanghai University of Engineering Science has developed a new approach to solve this specific problem, creating a system they call SQCAFusion. Instead of waiting until the end of the process to adjust for lighting conditions, their method asks the computer to understand the environment right at the very beginning. Imagine a translator who, before reading a book, first checks the time of day to decide whether to use formal or casual language; similarly, this new system first analyzes the scene to determine if it is day or night. It then uses this knowledge to actively guide the merging process from the start, rather than just applying a fix at the end. The researchers found that by injecting this "scene awareness" directly into the early stages of feature extraction, the computer can dynamically decide how much weight to give to the visible image versus the thermal image. If the visible image is dark and noisy, the system learns to trust the thermal data more, while still preserving the sharp edges of the visible world. If the scene is bright, it prioritizes the rich textures and colors of the standard camera.

The core of this innovation is a mechanism that acts as a dynamic filter. The system generates a set of digital tokens that represent the lighting conditions of the scene. These tokens then act as a query, asking the computer to look at the image features and decide which parts are important and which are just noise. This happens across multiple scales, meaning the system checks both the broad shapes of objects and the tiny details of textures simultaneously. By doing this, the computer can suppress the grainy noise that often plagues low-light photos without blurring the important targets, such as a pedestrian or a vehicle. The researchers also introduced a clever training strategy to handle the difficulty of learning from bad data. When the computer is trained on very dark images, the noise can sometimes trick the learning process into thinking that the grainy static is actually a real edge or object. To prevent this, the system was taught to automatically lower its expectations for the quality of the visible image when it detects a dark scene. This allows the model to ignore the noise while still faithfully capturing the true structure of the scene.

The results of this new framework were tested against a wide range of existing methods using several different datasets, including urban street scenes, military camouflage scenarios, and complex road environments. In standard tests on a dataset of over a thousand image pairs, the new method outperformed all previous state-of-the-art algorithms in key measures of information retention and visual clarity. It successfully preserved more detail and created images with higher contrast than its competitors. Perhaps more impressively, the system demonstrated a remarkable ability to generalize. When the researchers tested it on completely new datasets it had never seen before—ranging from high-definition road traffic scenes to single-channel military thermal images—the model did not need to be retrained or adjusted. It maintained its high performance, producing clear, sharp images that kept the thermal targets distinct while retaining the natural look of the background. In low-light conditions where other methods produced blurry halos or lost details entirely, this new approach kept the edges of objects crisp and the backgrounds clean.

The study confirms that the key to better image fusion lies not just in having powerful hardware, but in giving the software the ability to understand the context in which it is working. By moving away from a rigid, one-size-fits-all approach and toward a dynamic system that adapts to the lighting conditions in real-time, the researchers have created a tool that is far more robust for real-world applications. The work suggests that future vision systems will need to be context-aware to handle the unpredictable nature of the physical world. While the current model relies on a global understanding of the scene, the researchers note that future iterations could focus on even finer details, potentially allowing for real-time adaptation in even more complex environments. For now, this method represents a significant step forward in making machine vision reliable in the dark, ensuring that critical details are not lost to the shadows.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →