← Latest papers
🤖 AI

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models

This paper systematically investigates the impact of evolving LLM backbones (LLaMA-1 to LLaMA-3) on Vision-Language Models by controlling for other variables, revealing that newer backbones do not universally improve performance but instead alter task-specific behaviors, such as solving different visual questions with better-calibrated confidence, while offering limited benefits for tasks relying primarily on visual understanding.

Original authors: Sameera Horawalavithana, Lauren Phillips, Ian Stewart, Sai Munikoti, Karl Pazdernik

Published 2026-04-14
📖 5 min read🧠 Deep dive

Original authors: Sameera Horawalavithana, Lauren Phillips, Ian Stewart, Sai Munikoti, Karl Pazdernik

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, well-read Brain (the Large Language Model, or LLM) and a pair of Eyes (the Vision Encoder). To build a robot that can look at a picture and tell you a story about it, you need to connect the Eyes to the Brain.

This paper is like a scientific experiment where the researchers kept the Eyes exactly the same, but they swapped out the Brain three times. They used an older, slightly less advanced brain (LLaMA-1), a medium one (LLaMA-2), and the newest, most sophisticated one (LLaMA-3).

The big question they asked was: "If we give our robot a smarter brain, does it automatically get better at looking at pictures and answering questions?"

Here is what they found, explained simply:

1. A Smarter Brain Doesn't Always Mean a Smarter Robot

You might think, "Newer brain = Better robot." But the researchers found that it depends entirely on the job.

  • The "Memory Test" Job (ScienceQA): Imagine a test where you have to memorize a textbook chapter and then answer questions based only on that text.
    • Result: The newer brains actually did worse here. Why? Because the older brains were really good at memorizing the specific patterns in the training data. The newer, smarter brains started overthinking or relying on their own "common sense" instead of just reading the text, which caused them to make mistakes on this specific type of rigid test.
  • The "Real World" Job (VQA-Scene): Imagine looking at a photo of a messy living room and answering, "Is there a cat on the sofa?"
    • Result: The newer brains did better. They were better at understanding the context and the visual clues, leading to more correct answers.

The Lesson: Upgrading the brain doesn't fix everything. If the job is just about memorizing facts, a newer brain might overcomplicate things. If the job requires real understanding, a newer brain shines.

2. They Are Solving Different Puzzles

The researchers noticed something weird. When they looked at which specific questions the robots got right, they realized the newer robots weren't just getting the same questions right plus a few more. They were getting a completely different set of questions right.

  • Analogy: Imagine two detectives looking at the same crime scene.
    • Detective A (Old Brain) is very confident. They say, "I know who did it!" and point to the butler. They are 100% sure, but sometimes they are wrong.
    • Detective B (New Brain) is more humble. They say, "Hmm, it could be the butler, or maybe the gardener." They are less sure of themselves, but they are actually right more often.
    • The new brain isn't just "better" at the old way of thinking; it's looking at the clues in a totally different way.

3. The "Eyes" Are Still the Bottleneck

For some jobs, like analyzing a complex geological map to find a specific earthquake location, the researchers found that upgrading the brain did almost nothing.

  • Analogy: Imagine you have a super-genius brain, but you are trying to read a book written in a language you don't know, or the text is so blurry you can't see the letters. No matter how smart your brain is, if your Eyes (the vision encoder) can't see the details, the brain can't help.
  • In these cases, the researchers found that the newer brains were just as stuck as the older ones. To fix this, you don't need a smarter brain; you need better glasses (a better vision encoder).

4. The New Brain Can Do Magic the Old One Couldn't

There was one specific task where the newest brain (LLaMA-3) did something the older ones completely failed at: Predicting exact GPS coordinates from a picture of a landscape.

  • The older brains would look at the map and say, "It's in California."
  • The new brain looked at the map and said, "It's at Latitude 34.05, Longitude -118.25."
  • Why? The new brain had been trained on more data that included numbers and maps, so it learned a new "skill" that the older brains simply didn't have. It wasn't just an improvement; it was a brand-new ability.

5. The New Brain is More "Honest"

Finally, the researchers looked at how confident the robots felt about their answers.

  • The Old Brains were like overconfident students who always raised their hands and shouted the answer, even when they were guessing. They were 100% sure, even when wrong.
  • The New Brain was more like a careful scholar. It said, "I think the answer is X, but I'm only 70% sure."
  • Surprisingly, being less confident actually helped the new brain be more accurate. It didn't force an answer; it weighed the possibilities carefully.

Summary

The paper teaches us that in the world of AI, bigger and newer isn't always "better" in a straight line.

  • If you need a robot to memorize facts, an older, simpler brain might actually work better.
  • If you need a robot to understand complex scenes, a newer brain helps.
  • If the robot can't "see" well, giving it a smarter brain won't help; you need better eyes.
  • And sometimes, a new brain doesn't just get better at the old tricks; it learns entirely new magic tricks that the old ones couldn't even imagine.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →