← Latest papers
💬 NLP

Re:Verse -- Can Your VLM Read a Manga?

This paper introduces Re:Verse, a novel evaluation framework that systematically reveals current Vision Language Models' inability to perform deep narrative reasoning and maintain temporal causality in sequential visual storytelling like manga, despite their proficiency in individual panel recognition.

Original authors: Aaditya Baranwal, Madhav Kataria, Naitik Agrawal, Yogesh S Rawat, Shruti Vyas

Published 2026-04-23
📖 5 min read🧠 Deep dive

Original authors: Aaditya Baranwal, Madhav Kataria, Naitik Agrawal, Yogesh S Rawat, Shruti Vyas

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a robot friend who is incredibly smart at looking at single pictures. If you show it a photo of a cat, it knows it's a cat. If you show it a photo of a person holding a cup, it knows they are drinking.

But what happens if you hand that robot a comic book?

This paper, titled "Re:Verse," is like a report card given to the world's smartest AI robots (called Vision-Language Models) to see if they can actually read a manga (Japanese comic) and understand the story, or if they are just guessing based on individual pictures.

Here is the breakdown of what the researchers found, using some everyday analogies.

1. The Problem: The "Snapshot" vs. The "Movie"

Current AI models are like tourists with a camera. They take a great photo of a single moment (a single comic panel) and can describe it perfectly. "I see a boy looking scared. I see a cat hitting him."

But a comic book is a movie made of still pictures. To understand the story, you have to connect the dots between the pictures. You need to know:

  • Who is speaking? (Is the cat talking, or is the boy thinking?)
  • What happened before this picture?
  • What will happen next?
  • Why is the boy angry in panel 5, but happy in panel 6?

The researchers found that while these AI robots are great at describing the "snapshot," they are terrible at watching the "movie." They lose the thread of the story almost immediately.

2. The Test: "Re:Verse"

To test this, the team created a new exam called Re:Verse.

  • The Source Material: They used the first arc of a famous manga called Re:Zero. Why? Because Re:Zero is tricky. It involves time loops, characters dying and coming back, and complex emotional shifts. It's the "hard mode" of comic reading.
  • The Setup: They took 308 pages of the comic and manually matched every single speech bubble and thought bubble with the original text of the novel. It's like having a perfect script next to the comic book.
  • The Students: They tested 8 different AI models (both open-source and big ones) to see how well they could handle this exam.

3. The Results: The AI Got Lost

The results were a bit embarrassing for the AI. Here is what happened in three key areas:

A. The "Who Said What?" Disaster (Character Grounding)

Imagine a classroom where 10 kids are talking at once. If you ask a human, "Who said 'I'm hungry'?", they can point to the kid.
The AI, however, is like a confused substitute teacher.

  • It could often find the text bubbles (it knew words were there).
  • But when asked who was speaking, it guessed wrong almost 100% of the time.
  • The Metaphor: It's like watching a play where the actors swap scripts mid-scene. The AI sees the words, but it has no idea which character is holding the microphone.

B. The "Amnesia" Effect (Story Synthesis)

When asked to summarize the story or write a new chapter based on the pictures, the AI produced repetitive, boring nonsense.

  • The Metaphor: Imagine asking a robot to tell you a joke. It tells you the setup, then forgets the punchline, then repeats the setup, then forgets the punchline again.
  • The AI couldn't keep track of characters. It would call the main character "Subaru" in one sentence and then forget his name entirely in the next. It couldn't maintain a consistent personality for the characters over a long story.

C. The "Time Travel" Failure (Temporal Reasoning)

This was the biggest shock. The researchers asked the AI: "Here are pages 1 through 5. What happens on page 6?"

  • The Metaphor: It's like showing someone the first five frames of a movie and asking them to guess the sixth.
  • The AI performed terribly. It couldn't predict what would happen next. Even stranger, it sometimes did worse when there were fewer missing pages to guess. It seems the AI treats every page as a brand new, unrelated event, rather than part of a continuous timeline.

4. Why Does This Matter?

You might ask, "So what? It's just a comic book."

The researchers argue that comics are the perfect test for "real" intelligence.

  • If an AI can only recognize a picture of a dog, that's simple pattern matching.
  • If an AI can understand a story where a character dies, comes back to life, feels guilty, and then fights a villain, that requires deep reasoning.

The paper concludes that current AI is like a very fast reader who has no memory. It can read a sentence, but it forgets the previous page instantly. It lacks the ability to "connect the dots" across time.

The Takeaway

The Re:Verse benchmark is a wake-up call. It tells us that while our AI is getting better at seeing, it is still very bad at understanding stories.

To build a true "Storyteller AI" that can help writers, summarize books, or understand movies, we need to teach these models how to remember the past and predict the future, not just describe the present. Until then, if you want to read a manga, you're still better off doing it yourself!

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →