← Latest papers
💻 computer science

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

This paper introduces S-OBI, a novel benchmark that synthesizes clear, standardized sentence-level Oracle Bone Inscription instances to evaluate Multimodal Large Language Models, revealing that current models struggle with structured inscription understanding due to a heavy reliance on character-level recognition and the propagation of visual perception errors.

Original authors: Ziqi Li, Zijian Chen, Tingzhu Chen, Guangtao Zhai

Published 2026-07-01
📖 4 min read☕ Coffee break read

Original authors: Ziqi Li, Zijian Chen, Tingzhu Chen, Guangtao Zhai

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to read ancient Chinese history. For a long time, researchers have been teaching these robots to recognize individual ancient symbols, kind of like teaching a child to recognize the letter "A" or the number "1" in isolation. They can point to a single carved symbol on a turtle shell and say, "That's a 'horse'."

But real history isn't just a pile of random letters; it's full sentences with stories, dates, and specific meanings. This paper, titled "Beyond Single Character," argues that just because a robot can recognize the letters doesn't mean it can read the story.

Here is a simple breakdown of what the researchers did and what they found:

The Problem: The "Lego" vs. The "Castle"

Think of Oracle Bone Inscriptions (the ancient writing) as a castle made of Lego bricks.

  • Previous studies only tested if the robot could identify a single red brick or a single blue brick.
  • This paper asks: "Can the robot look at the whole castle and tell you who built it, why they built it, and what the story is?"

The researchers found that current AI models (called Multimodal Large Language Models, or MLLMs) are great at spotting the individual bricks but terrible at understanding the castle. They might know what a "horse" symbol looks like, but they can't figure out if the sentence is about a king hunting a horse or a horse running away.

The Solution: Building a "Clean" Test (S-OBI)

The ancient records are old, cracked, and dirty (like a muddy, torn-up photograph). To test the AI fairly, the researchers created a new benchmark called S-OBI.

Instead of using the messy, original photos, they acted like digital restorers:

  1. They took 95 ancient sentences that experts had already translated and understood.
  2. They took the messy, blurry parts of the ancient carvings and replaced them with crystal-clear, perfect digital versions of those same symbols.
  3. They kept the sentence structure exactly the same but made the visual input "clean."

This is like taking a muddy, torn-up letter from the 1800s, scanning it, and then re-typing it on a clean piece of paper with perfect font, so the AI isn't distracted by the dirt and can focus purely on understanding the meaning of the sentence.

The Test: Three Levels of Difficulty

They designed three types of puzzles to see how smart the AI really is:

  1. The "Gist" Test (Semantic Matching): "Here is an ancient sentence. Which of these four modern summaries describes it best?"
    • Result: The AI was okay at this. It could guess the general vibe.
  2. The "Structure" Test (Slot Extraction): "Here is a sentence. Please fill in a form with the Date, the Person, the Action, and the Object."
    • Result: The AI struggled. It couldn't organize the information correctly.
  3. The "Detective" Test (Contextual Reasoning): "Here is a sentence where we erased one word. Based on the other words, what was the missing word?" OR "Here is a sentence where we shuffled the words. Put them back in the right order."
    • Result: The AI failed miserably. It couldn't use the surrounding clues to figure out the missing piece.

The Big Discovery

The most surprising finding is that the AI is still stuck on the "single character" level.

The researchers found that if the AI makes a tiny mistake recognizing one small symbol, that error ripples through the whole sentence, causing the AI to guess the wrong meaning for the entire story. It's like if you misread one word in a sentence, you might misunderstand the whole paragraph.

Even the smartest AI models tested (like GPT-4o and Qwen) scored very low overall (around 33% correct). They could sometimes guess the right order of the words (like knowing the sentence starts with a date), but they couldn't actually understand the story the sentence was telling.

The Conclusion

The paper concludes that while AI is getting better at recognizing ancient symbols, it hasn't yet learned how to "read" them as a coherent language. It's like having a dictionary but not knowing how to write a poem. To truly understand these ancient records, AI needs to move beyond just identifying individual bricks and learn how to build (and understand) the whole castle.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →