MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
This paper introduces MME-VideoOCR, a comprehensive benchmark featuring 25 tasks across 44 video scenarios to evaluate Optical Character Recognition capabilities in Multimodal Large Language Models, revealing that current state-of-the-art models struggle with holistic video comprehension, spatio-temporal reasoning, and cross-frame integration despite achieving only 73.7% accuracy even with top-performing systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet hum of a modern computer, a new kind of intelligence is learning to see. These systems, known as multimodal large language models, are designed to process not just words, but also images and videos, blending them into a single stream of understanding. They have become remarkably skilled at reading text from static pictures, turning a photograph of a street sign or a document into digital words with high precision. This ability has opened doors for machines to navigate the world, read instructions, and organize information. However, the real world is rarely still. It moves, shifts, and changes. When these same intelligent systems are asked to read text from a moving video, the task becomes far more difficult. The text might blur as a car speeds by, appear only for a split second, or be scattered across different moments in time. Understanding text in motion requires more than just seeing; it demands remembering, connecting, and reasoning across a sequence of events.
A team of researchers has now taken a close look at how well these advanced systems handle text in the dynamic world of video. They built a new testing ground called MME-VideoOCR, a comprehensive collection of video clips designed to challenge machines in the specific ways that human eyes and minds struggle with moving text. This benchmark does not simply ask a computer to read a word; it asks the computer to find a specific number on a racing car, track a license plate as a vehicle changes lanes, or piece together a sentence that is broken up and scattered across several seconds of footage. The researchers gathered 1,464 videos, ranging from short clips to longer sequences, covering forty-four different real-world scenarios like driving, cooking, watching news, and attending lectures. For each video, they created two thousand carefully crafted questions and answers, all verified by human experts to ensure accuracy and fairness.
When they put eighteen of the most advanced video-reading models to the test, the results revealed a significant gap between current capabilities and the demands of the real world. Even the best-performing system, a powerful model known as Gemini-2.5 Pro, managed to answer correctly only 73.7 percent of the time. While this might seem like a high score, the researchers found that the models failed dramatically on tasks that required them to look at the whole video and connect information across time. For instance, when asked to recognize a letter formed by the path of a moving object, or to reconstruct a word from letters that appeared in random order across different frames, the top models scored near zero. They could read a sign clearly visible in a single frame, but they often lost the thread when the information was spread out.
The study also uncovered a peculiar habit in how these machines think. When faced with blurry or misspelled text, the models frequently ignored what was actually on the screen and instead guessed what the text should say based on common language patterns. If a sign in a video clearly read "OFF COURS" with a missing letter, the model would often confidently report it as "OFF COURSE," prioritizing its internal knowledge of how words are usually spelled over the visual evidence. This tendency to rely on what it expects to see, rather than what is actually there, suggests that these systems are still heavily influenced by their training on written language, sometimes at the expense of visual accuracy.
Furthermore, the researchers discovered that the quality of the video input matters immensely. When they fed the models lower-resolution images or fewer frames from the video, the performance dropped sharply. The machines needed high-definition clarity and a sufficient number of moments in time to piece together the story. This finding highlights that simply making a model larger or smarter is not enough; the way it sees the world must be sharp and continuous. The study concludes that while these artificial intelligences have made great strides in static image reading, they still lack the ability to fully comprehend the fluid, fragmented, and often messy nature of text in motion. The path forward requires not just better algorithms, but a deeper integration of visual memory and a resistance to the urge to guess based on habit.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.