WMS-Net:Wavelet-Mamba Synergistic Network for Joint Semantic Segmentation and Height Estimation of Remote Sensing Images
This paper proposes WMS-Net, a Wavelet-Mamba Synergistic Network that leverages multi-level Haar wavelet transforms, a Mamba-based Direction-Aware module, and a Cross-Task Attention Interaction mechanism to effectively address noise and feature entanglement, thereby achieving state-of-the-art joint semantic segmentation and height estimation from single RGB remote sensing images.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of remote sensing, where satellites and aircraft capture the Earth from above, a single photograph holds two distinct stories waiting to be told. One story is about what things are: a road, a tree, a building, or a car. This is known as semantic segmentation, a process that labels every pixel in an image with a specific category. The second story is about how high those things are. This is height estimation, which turns a flat picture into a three-dimensional map of the terrain, revealing the towering roofs of skyscrapers or the gentle slopes of a hill. For decades, computer scientists have tried to solve these two problems separately, training different algorithms to recognize objects and different ones to measure elevation. However, these two tasks are deeply connected; the shape of a building helps define its height, and knowing the height of a roof helps distinguish it from the ground. The challenge has been that when computers try to learn both at once, the noise and complexity of real-world images often cause the two lessons to get in each other's way, leading to blurry edges or inaccurate measurements.
A team of researchers at Xi'an University of Architecture and Technology has developed a new approach to untangle this relationship, creating a system they call WMS-Net. Instead of forcing a single algorithm to juggle both tasks simultaneously, they designed a network that splits the visual information into different layers of detail before recombining them. Imagine looking at a photograph and separating the sharp, crisp outlines of a building from the smooth, broad colors of the sky and ground. The researchers realized that the task of identifying what an object is relies heavily on those sharp edges and fine textures, while the task of measuring height depends more on the smooth, continuous shapes and overall contours. To exploit this difference, their new system uses a mathematical tool called a wavelet transform to act as a filter. This tool separates the image data into high-frequency components, which contain the fine details and edges, and low-frequency components, which hold the broader structural information.
Once the data is separated, the system directs the high-frequency details to the part of the network responsible for identifying objects, ensuring that the edges of buildings and trees are drawn with precision. Simultaneously, it sends the low-frequency, smooth information to the part of the network calculating height, allowing it to understand the overall shape and elevation without being distracted by noisy textures. This separation prevents the two tasks from confusing one another. However, simply splitting the data is not enough; the system also needs to understand how objects relate to one another across the entire image, such as how a long road stretches across a city or how a cluster of trees forms a continuous canopy. To achieve this, the researchers incorporated a component based on a modern architecture known as Mamba. This part of the system acts like a long-range scanner, capable of connecting distant parts of the image with very little computational effort, allowing it to model the global structure of the scene efficiently.
The final piece of the puzzle is a mechanism that allows the two separate streams of information—the detailed object labels and the smooth height map—to talk to each other. The researchers built a cross-attention module that lets the height information guide the object recognition, ensuring that a building's outline remains consistent with its measured elevation, and vice versa. This creates a feedback loop where the two tasks refine each other, correcting mistakes that might occur if they were working in isolation. When tested on two major datasets of aerial images from the cities of Vaihingen and Potsdam in Germany, this new system demonstrated superior performance compared to existing methods. On the Vaihingen dataset, the system achieved an accuracy score of 83.2% for identifying objects and a very low error rate for height measurements. On the larger, higher-resolution Potsdam dataset, it reached 86.8% accuracy for object identification and maintained a similarly low error rate for height.
The results suggest that by respecting the different ways humans and machines perceive detail versus structure, and by allowing these different views to collaborate rather than compete, it is possible to create a much more reliable tool for mapping the world. The system does not just produce a list of labels or a rough height map; it produces a coherent, three-dimensional understanding of the scene where the boundaries of a building align perfectly with its measured height. This approach offers a new framework for how computers can learn to see the world in multiple dimensions at once, turning flat photographs into rich, actionable data for urban planning, disaster monitoring, and smart city modeling. The success of this method lies not in making the computer smarter in a general sense, but in giving it a clearer way to organize the visual information it receives, separating the noise from the signal and the edges from the shapes before asking it to make sense of the whole.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.