← Latest papers
🧬 biology

Real-time Reconstruction of Human Visual Perception from fMRI

This paper presents a real-time compatible adaptation of the state-of-the-art MindEye2 pipeline using the RT-Cloud platform, demonstrating that reliable, fine-grained reconstruction of perceived natural images from fMRI data is feasible within seconds and paving the way for advanced brain-computer interfaces.

Original authors: Rishab S. Iyer, Jiaxin Cindy Tu, Cesar Kadir Torrico Villanueva, Anish Mahishi, Ross P. Kempner, Jacob S. Prince, Ernest W. Lo, Akash Bhowmick, Hritik Arasu, Amaar Chughtai, Elizabeth A. McDevitt, Pau
Published 2026-07-28
📖 7 min read🧠 Deep dive

Original authors: Rishab S. Iyer, Jiaxin Cindy Tu, Cesar Kadir Torrico Villanueva, Anish Mahishi, Ross P. Kempner, Jacob S. Prince, Ernest W. Lo, Akash Bhowmick, Hritik Arasu, Amaar Chughtai, Elizabeth A. McDevitt, Paul S. Scotti, Kenneth A. Norman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you could peek inside someone's mind and see exactly what picture they are looking at, just by watching the tiny ripples of blood flow in their brain. This is the dream of cognitive neuroscience: turning brain activity into a window on our thoughts. For a long time, scientists have used a tool called fMRI (functional magnetic resonance imaging) to do this. Think of fMRI like a super-slow, high-definition movie camera for the brain. It doesn't take pictures of neurons firing directly; instead, it tracks where oxygen-rich blood is flowing, which is a bit like seeing which parts of a city are busiest by watching where the traffic lights are glowing. The catch is that this "traffic" moves slowly, taking several seconds to peak after a thought or image appears.

In recent years, computers have gotten incredibly smart at guessing what a person is seeing based on these slow blood-flow patterns. Using massive AI models, scientists can now reconstruct images a person has seen with startling accuracy. However, there's a big problem: these smart computers usually need hours or even days to crunch the numbers. It's like having a genius chef who can cook a perfect meal, but they need a whole day to chop the vegetables. This makes it impossible to use them for "real-time" experiments, where you want to see the result instantly while the person is still in the scanner. If you could make this process fast enough, you could create a "brain-computer interface" where a person could control a screen with their thoughts, or get instant feedback on what their brain is focusing on.

This paper is about taking that slow, genius chef and teaching them to cook a meal in under 15 seconds. The researchers, led by Rishab Iyer and colleagues, successfully adapted a powerful, complex AI system called MindEye2 to work in real-time. They managed to decode what a person was seeing from an fMRI scan and reconstruct the image on a screen just seconds after the person looked at it. They didn't just do this in a perfect, slow-motion lab setting; they tested it on a standard 3 Tesla MRI scanner (the kind found in most hospitals, not just fancy research labs) and showed that it works even with just one hour of training data from a new person. While the images aren't quite as perfect as the slow, overnight versions, they are clear enough to recognize the object, proving that we can finally start "reading" minds as they happen, not just after the fact.

The "Mind-Reading" Machine Gets a Speed Boost

For years, the best AI systems for reading minds from brain scans have been like a slow-motion replay. You'd show a picture to a person in an MRI machine, wait for them to finish the experiment, and then spend hours or even days processing the data to guess what they saw. The paper describes a new approach that speeds this up dramatically, turning a "next-day" result into a "next-second" result.

The team used a system called MindEye2, which is a giant AI model trained to understand the connection between brain activity and images. Think of MindEye2 as a translator that speaks two languages: "Brain-Flow" (the fMRI data) and "Image-Description" (a digital map of what things look like). In the past, this translator needed a lot of time to think. The researchers wanted to see if they could make it think fast enough to work while the person was still in the machine.

To do this, they built a special pipeline using a cloud-based platform called RT-Cloud. This is like a high-speed delivery service that takes the brain data the moment it's recorded, processes it instantly, and sends the result back. They tested this with a participant looking at natural images, like pictures of animals or places.

The Results: Fast, But Not Perfect

The paper presents a fascinating trade-off between speed and quality. They tested three different speeds:

  1. The "Fast" Mode: This is the real-time magic. The system waited about 7.9 seconds after the image appeared (to let the brain's blood flow response peak) and then did the decoding. The total time from seeing the image to getting the reconstruction was about 14.5 seconds.
  2. The "Slow" Mode: They waited about 29 seconds to gather more data, which improved the quality slightly.
  3. The "End-of-Run" Mode: This was the traditional approach, waiting until the whole scanning session was over (about 5 minutes) to analyze everything.

The results showed that even the "Fast" mode worked surprisingly well. When asked to pick the correct image out of a pool of 50 candidates, the fast system got it right about 36% of the time (which is much better than the 2% you'd get by guessing randomly). When the pool was smaller (just 2 options), the accuracy jumped to 90.6%. The reconstructed images were a bit blurry compared to the slow, overnight versions, but you could still clearly tell if the person was looking at a dog, a car, or a landscape.

Why This Matters (and What It Doesn't Do Yet)

The paper is careful to say this is a "proof-of-concept." It proves that it is possible to do this in real-time, but it doesn't mean we have a perfect mind-reading device yet. The authors explicitly note that the quality drops when you go faster. The "Fast" mode is not as accurate as the "Offline" mode that takes days. They also point out that the current system relies on a specific type of AI architecture and that the speed is limited by how fast blood flows in the brain (the hemodynamic lag), which is a biological fact we can't change.

However, the implications are huge. Because the system works in under 15 seconds, it opens the door for neurofeedback. Imagine a patient with depression who could see, in real-time, how their brain reacts to a sad picture. If their brain is focusing too much on the negative parts, the system could show them a reconstruction that highlights that, helping them learn to change their focus. The paper suggests that delays of 10 to 30 seconds are actually fast enough for the brain to learn from this feedback, based on previous studies.

The researchers also showed that you don't need a super-expensive, rare 7 Tesla scanner to do this; a standard 3 Tesla scanner (found in thousands of hospitals) works just fine, as long as you fine-tune the AI with about an hour of data from the new person.

The Bottom Line

This paper is a major step forward in turning "brain-reading" from a slow, post-mortem analysis into a live, interactive experience. By adapting a heavy-duty AI to run in real-time, the team showed that we can reconstruct what a person is seeing within 14.5 seconds. While the images aren't crystal clear yet, they are good enough to recognize objects, and the speed is fast enough to potentially help people control their own brain states or communicate with machines. It's like taking a slow, high-definition movie camera and teaching it to stream live video. The picture is a little grainy, but for the first time, we can see the show as it happens.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →