← Latest papers
💻 computer science

Conv-Guided Mamba Vision Transformer and SPP-DenseNet121 based Hybrid Multi-Feature Fusion and Query Expansion in Content-Based Medical Image Retrieval

This paper proposes a novel Content-Based Medical Image Retrieval framework that integrates a Conv-Guided Mamba Vision Transformer with SPP-DenseNet121 for hybrid multi-feature fusion and employs query expansion to achieve high retrieval accuracy on endoscopic bladder tissue and gastrointestinal datasets.

Original authors: Mathana Gopal Arulsamy, Pallikonda Rajasekaran Murugan, Md. Jakir Hossen, Wai Kit Wong, Poh Kiat Ng, Gomathy Nayagam M

Published 2026-07-17
📖 4 min read☕ Coffee break read

Original authors: Mathana Gopal Arulsamy, Pallikonda Rajasekaran Murugan, Md. Jakir Hossen, Wai Kit Wong, Poh Kiat Ng, Gomathy Nayagam M

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking through a massive, endless library where every single book is a photograph of the inside of a human body. If you were a doctor trying to find a picture of a specific type of bladder tumor to help a patient, you wouldn't want to flip through millions of pages one by one. You'd want a super-smart librarian who could instantly grab the exact photos you need. This is the world of Content-Based Medical Image Retrieval (CBMIR). Instead of searching by text labels like "cancer" or "polyp," this system searches by looking at the actual picture content—the colors, shapes, and textures. The challenge is that medical images are tricky; two pictures might look almost identical but mean very different things to a doctor, or two pictures of the same disease might look wildly different depending on how the camera was held or the lighting. The goal is to build a digital detective that understands both the tiny details and the big picture to find the right medical images quickly and accurately.

In this study, the researchers built a new kind of digital detective called a Hybrid Multi-Feature Fusion and Query Expansion framework. Think of it as a team of two expert detectives working together to solve a case. The first detective is a specialist in local details, using a tool called SPP-DenseNet121. This detective is great at spotting tiny clues like the texture of a tissue or the edge of a lesion, no matter how big or small the spot is. The second detective is a master of global context, using a new, efficient tool called Conv-Guided Mamba Vision Transformer (CG-MViT). This detective looks at the whole scene to understand how different parts of the image relate to each other over long distances, kind of like understanding the layout of a whole room rather than just a single brick.

Usually, when you combine two different sets of clues, you might end up with a messy pile of redundant information. To fix this, the team used a special math trick called Discriminative Multiple Canonical Correlation Analysis (DMCCA). Imagine this as a super-organized filing system that takes the notes from both detectives, removes the duplicates, and merges them into one perfect, crystal-clear report. But the team didn't stop there. They knew that sometimes the first guess isn't perfect, so they added a "second-guess" step called Pseudo-Relevance Feedback (PRF). This is like asking the computer, "Hey, the top 10 pictures you found look good, right? Let's assume they are all correct and use them to refine our search." This helps the system understand the doctor's intent better and find even more relevant images.

The researchers tested their new detective team on two real-world medical photo libraries: the Kvasir dataset (which contains images of the gastrointestinal tract) and the Endoscopic Bladder Tissue dataset. The results were impressive. On the Kvasir dataset, the system found the right images with a high degree of accuracy, scoring a 97.72% success rate when looking at the top 5 results, and still maintaining a strong 91.02% even when looking at the top 100 results. On the more challenging bladder tissue dataset, it achieved 92.60% accuracy for the top 5 results and 66.82% for the top 100.

To make sure their new system was actually the hero and not just lucky, the researchers ran a series of "ablation studies." This is like taking apart a car engine to see which part makes it go fastest. They tested the system with just the first detective, then just the second, then both without the filing system, and finally the full team. They found that every single part contributed to the success; the full team with the filing system and the "second-guess" step performed the best. The study suggests that by combining local and global views, organizing the data smartly, and letting the system learn from its own initial guesses, medical image retrieval can become much more reliable. This could help doctors find similar past cases faster, aiding in diagnosis and treatment planning, though the authors note that more testing on larger and different types of medical images is needed to confirm how well this works in every hospital setting.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →