← Latest papers
💻 computer science

Uncovering neural decisions: How textual interpretation bridges AI and clinical reasoning

The paper introduces CRysTal, a neuron-level post-hoc explanation framework that maps AI model activations to standardized clinical terminology, demonstrating superior alignment with radiology reports, enhanced diagnostic reliability, and improved efficiency for physicians compared to existing explanation methods across multiple medical imaging tasks.

Original authors: Hui Lu

Published 2026-09-08
📖 6 min read🧠 Deep dive

Original authors: Hui Lu

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of modern medicine, artificial intelligence has become a powerful partner for doctors, capable of scanning images of the human body with a speed and precision that often surpasses human vision. These computer systems can look at a picture of a tumor and decide whether it is likely to be harmless or dangerous. However, a significant barrier stands between these powerful tools and their full acceptance in hospitals: the "black box" problem. While a computer can give a correct answer, it often cannot explain how it reached that conclusion. It might point to a blurry patch of an image, but it cannot tell a doctor why that patch matters. In medicine, knowing the reason is just as important as the result. Doctors rely on specific, standardized signs—like the shape of a lump, the sharpness of its edges, or how sound waves bounce off it—to make life-or-death decisions. If a computer cannot speak the language of these signs, doctors cannot trust its advice, especially when the computer and the human doctor disagree.

A researcher at Shanghai Jiao Tong University and other institutions has developed a new method called CRysTal to solve this problem of trust. Their work focuses on breast ultrasound imaging, a common tool used to examine breast tissue. The researcher wanted to see if they could translate the internal, invisible calculations of an artificial intelligence model into clear, written descriptions that match the standard medical language doctors use every day. Instead of just showing a heatmap of where the computer looked, CRysTal acts as a translator. It takes the raw electrical signals from the computer's brain and maps them to a structured list of clinical features, such as "irregular shape," "angular margin," or "microcalcifications." This process creates a chain of evidence that a doctor can read and verify, turning a mysterious prediction into a transparent report.

To test this system, the researcher trained several different types of artificial intelligence models to distinguish between benign and malignant breast lesions using thousands of ultrasound images. Once these models were trained and their internal settings were locked, the researcher applied CRysTal to interpret them. The system works by analyzing how specific parts of the computer's brain, known as neurons, fire when it sees an image. It then uses a mathematical process to figure out which clinical concepts these neurons are responding to. For example, if a group of neurons lights up when the computer sees a jagged edge, CRysTal learns to label that activity as "angular margin." By combining these individual neuron signals, the system builds a complete, structured explanation for every single diagnosis it makes.

The results showed that this new approach was far more accurate than previous methods at explaining what the computer was thinking. When the researcher compared the explanations generated by CRysTal against the actual reports written by human radiologists, they found a much higher level of agreement. The system correctly identified key features like irregular shapes and specific types of margins far more often than other explanation tools. In fact, across the different computer models they tested, CRysTal improved the accuracy of these feature descriptions by between 3.4% and 12.2% compared to the next best method. It also did a better job of capturing the overall direction of the diagnosis, correctly identifying whether the evidence pointed toward a benign or malignant condition. This was a crucial finding because it proved that the computer was not just guessing; it was actually using the same visual clues that human experts rely on.

The researcher also watched how the computer learned over time. They tracked the explanations as the models were being trained from scratch. In the early stages, the computer's reasoning was vague and based on general visual patterns. But as the models got better at making correct diagnoses, their internal reasoning became increasingly aligned with standard medical concepts. The computer began to focus on the specific signs doctors look for, such as the orientation of the lesion or the presence of shadowing behind it. This shift suggests that when an artificial intelligence model learns to perform well, it naturally starts to think more like a human expert, and CRysTal was able to reveal this evolution.

To see if real doctors found these explanations useful, the researcher conducted a study with twenty ultrasound physicians, ranging from junior residents to senior experts. The doctors were asked to read the structured explanations generated by CRysTal and rate them on accuracy, readability, and how well they matched medical guidelines. The doctors gave the explanations extremely high scores, averaging close to 97 out of 100, with no significant difference between the experience levels of the doctors. This indicated that the language used by the system was clear and professional enough for anyone in the field to understand.

The study went a step further to see if these explanations could actually help doctors make better decisions. In a simulated workflow, doctors first looked at images alone, then looked at them with just a computer score, and finally looked at them with the computer score plus the CRysTal explanation. When the computer made a mistake, the doctors were often misled by the score alone, but the structured explanation helped them catch the error. With CRysTal, junior doctors improved their accuracy significantly when the computer was wrong, and senior experts also performed better. Furthermore, the structured explanations helped doctors work faster, reducing their reading time by about 10% for junior doctors and nearly 5% for senior experts. This suggests that clear, text-based evidence allows doctors to process information more efficiently than when they are forced to interpret vague visual heatmaps.

The researcher also tested whether this method could work on other types of medical imaging. They applied CRysTal to datasets involving thyroid ultrasound and skin lesion imaging, using the specific medical terminology systems for those areas. The system successfully adapted to these new tasks, generating explanations that were consistent with the clinical findings for thyroid nodules and skin spots. This flexibility suggests that the approach is not limited to breast cancer but could be a general tool for making artificial intelligence transparent across different areas of medicine.

Ultimately, this work demonstrates that it is possible to bridge the gap between the complex mathematics of artificial intelligence and the practical language of clinical medicine. By translating the internal signals of a computer into a structured list of medical facts, CRysTal provides a way to audit the decisions of AI systems. It does not change how the computer makes its prediction, but it opens a window into the process, allowing doctors to verify the evidence behind every diagnosis. This transparency is a vital step toward building a future where artificial intelligence can be a trusted, collaborative partner in the clinic, helping to ensure that every decision is safe, reliable, and understandable.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →