← Latest papers
💻 computer science

Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI

The paper proposes PRISM, a novel framework that projects fMRI signals into a structured text space and employs object-centric diffusion with attribute-relationship search to achieve state-of-the-art visual stimulus reconstruction by better aligning with the compositional nature of neural activity.

Original authors: Zheng Huang, Enpei Zhang, Weikang Qiu, Yinghao Cai, Carl Yang, Elynn Chen, Xiang Zhang, Rex Ying, Dawei Zhou, Yujun Yan

Published 2026-05-15
📖 5 min read🧠 Deep dive

Original authors: Zheng Huang, Enpei Zhang, Weikang Qiu, Yinghao Cai, Carl Yang, Elynn Chen, Xiang Zhang, Rex Ying, Dawei Zhou, Yujun Yan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine your brain is a busy library where every time you see an image, it doesn't just "store" the picture. Instead, it writes a very specific, complex story about what it sees. For years, scientists have tried to read these stories (from fMRI brain scans) and turn them back into the original pictures you saw.

The paper you're asking about, PRISM, is like a new, smarter translator that finally cracked the code on how to do this better. Here is the simple breakdown of what they found and built:

1. The Big Surprise: The Brain Speaks "Text," Not "Pictures"

Most previous attempts to rebuild images from brain scans tried to translate the brain's signal directly into a "picture language" (like a digital photo file). They assumed that because you saw a picture, your brain must be storing it as a picture.

The PRISM team discovered this was wrong.
They found that the brain's activity actually looks much more like text (like the sentences in a book) than it does like a photo or a video.

  • The Analogy: Imagine trying to translate a foreign language. Previous methods tried to translate the brain's "language" into another visual language (like translating a French movie directly into a German movie). PRISM realized the brain is actually speaking English (text). So, the best way to understand it is to translate it into English first, and then use that English to describe the picture.

2. The Problem with "Blurry" Descriptions

Even when scientists used text, they often made a mistake. They would write a single, long sentence describing the whole image, like: "A gray tiger-striped cat is sitting on a bench."
When a computer tries to draw this, it often gets confused. It might draw a "gray tiger" instead of a "cat," or mix up the colors. This is called an "attribute binding error"—the computer knows the words but forgets which word belongs to which object.

The Human Brain doesn't do this.
When you look at a scene, your brain doesn't see one giant blob. It sees distinct items: "There is a cat. It is gray. It has stripes. It is on a bench." It builds the scene piece by piece.

3. The Solution: PRISM (The "Lego" Builder)

The authors built a new system called PRISM (Projects fMRI sIgnals into a Structured text space). Think of it as a master builder who uses Lego bricks instead of a single giant mold.

Here is how PRISM works, step-by-step:

  • Step 1: The Translator (fMRI to Text)
    The system takes the brain scan and translates it into a structured text list. Instead of one long paragraph, it breaks the image down into specific parts:

    • Object 1: A car.
    • Object 2: An elephant.
    • Relationship: The car is following the elephant.
    • Attributes: The car is white; the elephant is gray.
  • Step 2: The Smart Search (Finding the Right Words)
    The system uses a "search engine" to figure out exactly which words matter most. It tested thousands of different ways to describe things and found that spatial words (like "left," "right," "on top of," "next to") are the most important for matching the brain's activity. It's like realizing that to describe a room to someone blindfolded, telling them where things are is more important than telling them what color the curtains are.

  • Step 3: The Lego Builder (Object-Centric Generation)
    Instead of asking an AI to draw the whole picture at once, PRISM asks it to draw the car first, then the elephant, and then puts them together in the right spots based on the text description.

    • Why this works: It stops the computer from getting confused. It ensures the car stays a car and the elephant stays an elephant, and they don't accidentally merge into a "car-elephant."

4. The Results

When they tested this on real people looking at images, PRISM did a much better job than previous methods.

  • The Score: It reduced the "perceptual loss" (how much the new image looks different from the original) by about 6%.
  • The Proof: If you showed the reconstructed images to a smart AI and asked, "What is in this picture?", the AI could answer correctly much more often with PRISM's images than with older methods.

Summary

The paper claims that to read the brain's visual thoughts, we shouldn't try to force the brain to speak in "pixels." Instead, we should listen to it as if it were writing a structured story about objects and their locations. By translating brain scans into this specific type of "object-based text" and then building the image piece-by-piece, we can see what the person was looking at with much greater clarity.

Note: The paper focuses strictly on the technical method of reconstructing static images from brain scans. It does not claim to have clinical applications, medical uses, or the ability to read thoughts in real-time for communication at this stage.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →