← Latest papers
🤖 AI

Concepts Learned Visually by Infants Can Contribute to Visual Learning and Understanding in AI Models

This paper demonstrates that incorporating early-acquired human-like visual concepts, such as animacy and goal attribution, into AI models significantly enhances their learning efficiency, accuracy, and generalization in predicting future events compared to standard deep network approaches.

Original authors: Shify Treger, Shimon Ullman

Published 2026-03-27
📖 5 min read🧠 Deep dive

Original authors: Shify Treger, Shimon Ullman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to understand the world. You have two different ways to do it.

Method A (The "Naive" Way): You throw a massive pile of raw video footage at the robot and say, "Figure out what's going on and guess what happens next." The robot has to learn everything from scratch: what a hand is, what a ball is, what it means to "want" something, and how things move. It's like trying to learn to drive a car by staring at a pile of engine parts and hoping you intuitively understand how the steering wheel works.

Method B (The "Cognitive" Way): You give the robot a few simple, foundational rules first. You say, "First, learn that some things move on their own (like people or dogs) and some things don't (like rocks or books). Second, learn that things that move on their own usually have a goal." Once the robot understands these basic concepts, you then show it the complex videos and ask it to guess what happens next.

This paper, written by researchers Shify Treger and Shimon Ullman, argues that Method B is much better. They show that when AI models learn like human babies do—by mastering simple concepts first and building on them—they become smarter, faster, and more efficient.

Here is a breakdown of their findings using some everyday analogies:

1. The "Baby Brain" vs. The "Data Dump"

Human babies are amazing learners. By the time they are a few months old, they can tell the difference between a living thing (like a hand) and a non-living thing (like a toy car). They also intuitively know that a living thing is trying to get something (a goal), while a non-living thing just follows physics (like rolling down a hill).

The researchers built two computer models to test this:

  • The "Naive" Model: This is the standard AI. It tries to learn the whole complex task (predicting the future) all at once, without any help.
  • The "Cognitive" Model: This model is given a "cheat sheet" of early human concepts. It learns to identify "animate" (living/moving on its own) vs. "inanimate" objects first. Then, it uses that knowledge to predict what will happen next.

The Result: The "Cognitive" model was like a student who studied the alphabet before trying to write a novel. It learned the task perfectly and much faster than the "Naive" model, even when they were given very little data. The Naive model struggled, often getting it wrong, even after seeing thousands of examples.

2. The "Magic Trick" of Generalization

Imagine you teach a child that a dog chases a ball.

  • The Naive Model is like a child who memorizes that specific dog chases that specific ball. If you swap the ball for a frisbee, or the dog for a cat, the child is confused.
  • The Cognitive Model is like a child who understands the concept of "chasing." If you show them a cat chasing a frisbee, they instantly get it.

The researchers tested this by changing the scenes. When they introduced new actors (like a new type of animal) or new objects, the "Cognitive" model adapted instantly. The "Naive" model got stuck. It was as if the Naive model was trying to memorize a map of a single city, while the Cognitive model learned the rules of navigation so it could find its way in any city.

3. The "Goal" Problem

The researchers also tested a specific puzzle called the "Woodward Task."

  • The Scene: A person reaches for a toy. Then, the toys switch places.
  • The Question: Where will the person reach next?
    • If it's a person (animate), they will reach for the toy (the goal), even if it moved.
    • If it's a rolling suitcase (inanimate), it will just keep rolling to the same spot where it stopped before, regardless of what's there.

The Shocking Discovery: The researchers tested the world's most advanced AI models (like the latest versions of GPT, Gemini, and Claude) on this task.

  • Humans: Got it right almost every time.
  • AI Models: Many of them got it wrong! Even though these models can write poetry and code, they failed to understand the simple difference between "I want that toy" and "I am just rolling to that spot."

The AI models were so focused on the visual patterns that they missed the meaning behind the movement. They didn't have the "early concept" of "goal" to guide them.

4. Why This Matters

The paper suggests that we are trying to build AI by throwing more data at it (like feeding a baby a mountain of food), but we aren't giving it the right "digestive system" to process it.

The Solution: Instead of just making AI bigger, we should make it smarter by teaching it the "ABCs" of human vision first.

  • Teach the AI to recognize Animacy (what moves on its own).
  • Teach it Goal Attribution (what things are trying to do).
  • Let it use those simple tools to solve complex problems.

The Big Picture Takeaway

Think of learning like building a house.

  • Current AI is trying to build the roof before the foundation is even poured. It's shaky, inefficient, and prone to collapse when the wind blows (new data).
  • The "Cognitive" approach says: "Let's lay the bricks of 'Animacy' and 'Goals' first." Once that foundation is solid, the rest of the house (complex understanding) goes up quickly, stands tall, and can handle storms.

The authors conclude that if we want AI to truly understand the world like humans do, we need to stop treating it like a blank slate and start teaching it the same simple, early lessons that human babies learn naturally.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →