← Latest papers
💬 NLP

Artificial Phantasia: Emergent Mental Imagery in Large Language Models

This study demonstrates that Large Language Models can outperform humans in a task traditionally requiring pictorial mental imagery by relying solely on propositional representations, thereby providing evidence for an emergent "artificial phantasia" that challenges the cognitive science view that visual imagery necessitates pictorial formats.

Original authors: Morgan McCarty, Jorge Morales

Published 2026-05-20
📖 5 min read🧠 Deep dive

Original authors: Morgan McCarty, Jorge Morales

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Can You "See" Without Eyes?

Imagine you are asked to close your eyes and picture a capital letter D. Now, imagine turning that D 90 degrees to the left. Finally, imagine attaching a capital letter J to the bottom center of that shape.

If you open your eyes, what do you see? Most humans would say, "It looks like an umbrella."

For decades, cognitive scientists have believed that to solve this kind of puzzle, your brain needs a special "mental screen" or a "mind's eye" where you can actually see and manipulate pictures. They thought you couldn't do it just by thinking about words; you needed a visual picture in your head.

This paper asks a wild question: Can a computer do this without ever having "eyes" or a "mind's eye"?

The Experiment: A Mental Gym for AI and Humans

The researchers created a gym workout for the brain. They took a classic puzzle (the letter transformations) and invented 48 brand-new, never-before-seen versions of it. Because these puzzles were made up specifically for this study, the computers couldn't have memorized the answers from their training data.

They asked two groups to solve these puzzles:

  1. 100 Humans: People who had to listen to the instructions (to avoid seeing the letters on a screen) and imagine the shapes in their heads.
  2. Large Language Models (LLMs): Advanced AI chatbots (like the latest versions of OpenAI's o3, GPT-5, and Google's Gemini 3 Pro). These models are trained on text, not on "seeing" pictures.

The Shocking Result: The AI "Saw" Better Than Humans

The results were surprising. The best AI models didn't just pass; they crushed the human average.

  • The Humans: Scored an average of 2.85 out of 5.
  • The Top AIs: Scored between 3.13 and 3.65.

The AI models solved these "visual" puzzles significantly better than the human participants. Even more interestingly, when the researchers forced the AI to actually draw the images at each step (using an image generator), the AI got worse at the task. This suggests that the AI wasn't relying on drawing pictures to solve the problem; it was doing something else entirely.

The "Magic Trick": How Did They Do It?

If the AI didn't use a "mind's eye" (a picture) and didn't use a drawing tool, how did it solve the puzzle?

The authors propose a concept they call "Artificial Phantasia."

Think of it like this:

  • The Human Way: You are an architect. You close your eyes and build a 3D model of a house in your mind. You walk around it, look at the windows, and then describe it.
  • The AI Way: You are a master translator. You never build a 3D model. Instead, you treat the instructions like a complex recipe written in code. You know that "Letter D rotated left" + "Letter J attached" = "Umbrella shape" purely through the logic of language.

The paper suggests these AI models are using propositional reasoning. They are manipulating words and concepts so precisely that they can figure out the final shape without ever "seeing" it. It's like solving a math problem by knowing the rules of algebra, rather than drawing the shapes on a piece of paper.

The "Reasoning" Factor

The researchers also tested what happens if they tell the AI to "think harder" (by giving it more time and "reasoning tokens" to process the steps).

  • Result: The more the AI was allowed to "think" and chain its reasoning together, the better it got.
  • Analogy: Imagine a detective solving a mystery. If you give them 5 minutes to look at the clues, they might guess. If you give them 5 hours to connect every clue logically, they solve the case perfectly. The AI solved the visual puzzle by connecting linguistic clues, not by looking at a picture.

What About "Blind" Humans?

The paper also mentions a group of people called aphantasics. These are humans who claim they cannot visualize images in their minds at all (they have no "mind's eye"). Surprisingly, studies show these people can still solve these visual puzzles almost as well as people who can visualize.

This paper adds fuel to that fire. It suggests that maybe you don't need a "mind's eye" to solve visual puzzles. You might just need a very strong "mind's logic." The AI, which has no mind's eye at all, proved that language and logic alone might be enough to do what we thought required a picture.

The Bottom Line

This study challenges the old idea that "visual thinking" requires a visual picture. It shows that advanced AI can perform tasks that look like "visual imagination" using only language and logic.

  • The Claim: AI has an emergent ability to "imagine" visually, but it's not doing it with pictures. It's doing it with words.
  • The Implication: We might need to rethink how we define "mental imagery." It might not be about seeing pictures in your head; it might just be about manipulating concepts logically.

The paper concludes that these AI models have discovered a new way to "see" the world—one that is purely linguistic, yet surprisingly effective at solving visual problems.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →