Quantification and object perception in Multimodal Large Language Models and human linguistic cognition
This paper investigates how Multimodal Large Language Models encode human-like quantification features—specifically quantifier scales, prototypicality, and approximate number biases—revealing that while "thinking" models excel at numerosity estimation and scaling, they still diverge significantly from human cognition in usage ranges and prototypicality across different languages and model types.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to look at a jar filled with red and blue marbles and describe how many of each color it sees. You ask, "Are there some red ones? Many? Most?"
To a human, this is easy. We have a built-in "number sense" that helps us guess the count, and we have a social understanding of what words like "some" or "most" actually mean in conversation. But for Artificial Intelligence (AI), specifically Multimodal Large Language Models (MLLMs)—the super-smart computers that can see images and read text—this task is surprisingly tricky.
This paper is like a detective story where researchers try to figure out why these AI models get stuck on simple counting and describing tasks, and whether they think like humans or just mimic us.
Here is the breakdown of their investigation using some everyday analogies:
1. The Three Suspects (The Research Questions)
The researchers wanted to solve three mysteries:
- Mystery A: Why do AI models struggle with words like "some" and "most"? Is it because they can't count, or because they don't understand the meaning of the words?
- Mystery B: Do "Thinking" models (AI that pauses to reason step-by-step) do better than "Instruct" models (AI that just answers quickly)?
- Mystery C: Does the language matter? Do they act differently in English, Spanish, or Greek?
2. The Experiment: The "Marble Jar" Test
The researchers showed humans and AI models pictures of black squares and white circles.
- The Human Task: Look at the picture, guess roughly what percentage is black, and pick the best word to describe it (e.g., "A few," "Some," "Many," "Most").
- The AI Task: Do the exact same thing.
It's like a game of "Guess the Crowd." If 60% of the crowd is wearing hats, a human might say "Many." If 90% are wearing hats, a human says "Most." The researchers wanted to see if the AI followed these same rules.
3. The Findings: What the AI Got Wrong
The "Counting" Problem (The Approximate Number System)
Humans have a natural "gut feeling" for numbers. If you see a jar that is 80% full, we know it's "most."
- The Result: The "Thinking" models (the ones that pause to reason) were actually pretty good at guessing the numbers. They were like a student who studied hard for a math test.
- The Problem: The "Instruct" models (the quick responders) often guessed wrong. They tended to think there were more black squares than there actually were. They were like someone guessing the number of jellybeans in a jar and wildly overestimating.
The "Word Meaning" Problem (The Scale of Words)
Humans organize words like a ladder:
A few (bottom rung) → Some → Many → Most (top rung).
If you say "Most," you imply "More than Many."The Result: The "Thinking" models mostly understood this ladder. They knew "Most" was higher than "Many."
The Problem: The "Instruct" models got the ladder mixed up. Sometimes they used "Most" for 60% and "Many" for 70%. It's like a child who thinks "Super" is a bigger word than "Mega" just because they heard it in a different context. They didn't understand the strict rules of the ladder.
The "Sweet Spot" Problem (Prototypicality)
This is the most interesting part. Humans have a "sweet spot" for words.
- If you ask a human, "When do you say 'Most'?" they will almost always say, "When it's like 90% or 95%."
- If you ask the AI, "When do you say 'Most'?" they often say, "When it's 55% or 60%."
The Analogy: Imagine a thermostat. Humans set "Most" to "Very Hot" (90 degrees). The AI sets "Most" to "Warm" (55 degrees). Even if the AI counts the squares correctly, it uses the wrong word for the temperature. It's like calling a lukewarm soup "boiling."
4. The Twist: Thinking vs. Instruct
The researchers found that the "Thinking" models (which do extra math and reasoning) were much closer to humans than the "Instruct" models.
- Thinking Models: Like a careful accountant. They count well and understand the ladder of words, but they still pick the "sweet spot" a bit too low.
- Instruct Models: Like a fast-talking salesperson. They guess the numbers wrong and mix up the word ladder completely.
5. The Language Lesson
The researchers tested this in six languages (English, Greek, Russian, Spanish, Italian, Catalan).
- The Surprise: They expected the AI to be best at English (since it's trained mostly on English). But the AI struggled just as much, or sometimes more, in other languages.
- The Conclusion: The AI doesn't seem to have one "universal brain" for numbers and words that works the same in every language. Instead, it seems to have different, slightly broken "dictionaries" for each language.
The Big Takeaway
This paper tells us that while AI is getting better at seeing and counting, it still doesn't truly understand the "flavor" of human language.
- Humans use words like "some" and "most" based on a mix of math, social rules, and gut feelings.
- AI uses them based on statistical patterns it saw in its training data. It's like a parrot that learned to say "Most" when it sees a lot of things, but it doesn't feel the difference between "a lot" and "almost all."
The "Thinking" models are a step in the right direction, but until AI can understand the subtle, fuzzy boundaries of human words, it will keep saying "Most" when we really mean "Many."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.