← Latest papers
🤖 AI

Seeking Information with RAG-Assistants: Does Model Size Matter in Human-AI Collaborations?

This study evaluates RAG-assisted chatbots in realistic multi-turn information-seeking scenarios, finding that while human-AI collaboration significantly improves performance regardless of underlying model size, user satisfaction and perceived usability remain largely unaffected by the model's scale.

Original authors: Lennard C. Froma, Tom Kouwenhoven, Maaike H. T. de Boer, Catholijn M. Jonker, Max J. van Duijn

Published 2026-05-06
📖 5 min read🧠 Deep dive

Original authors: Lennard C. Froma, Tom Kouwenhoven, Maaike H. T. de Boer, Catholijn M. Jonker, Max J. van Duijn

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to solve a complex puzzle, but the instructions are hidden inside a massive, 400-page manual written in technical jargon. You have two tools to help you: a smart assistant (an AI chatbot) and the manual itself.

This paper is like a report card on how well different sizes of these "smart assistants" work when humans team up with them to solve these puzzles. The researchers wanted to know: Does having a "bigger brain" (a larger AI model) actually make the human-AI team work better, or is a smaller, cheaper brain just as good when a human is in the loop?

Here is the breakdown of their experiment and findings, using simple analogies:

The Setup: Three Types of Assistants

The researchers created three versions of the AI assistant, each with a different "brain size" (model size):

  1. The Compact Brain (3B): Small, fast, and efficient. Think of this as a very smart intern who knows the basics but might miss subtle details.
  2. The Mid-Size Brain (8B): A step up. Think of a junior manager who is quite capable.
  3. The Giant Brain (70B): A massive, powerful model. Think of a senior expert with a huge library of knowledge.

The Twist: None of these assistants were allowed to just "guess" from their memory. They were all connected to the 400-page manual (a technique called RAG, or "Retrieval-Augmented Generation"). This is like giving the assistant a direct phone line to the library so they can look up facts instantly, rather than relying on what they memorized in school.

The Experiment: Humans vs. Machines

The researchers hired 112 people to act as "pilots" (even though they weren't real pilots) and asked them to answer specific questions about the manual. They were split into three groups:

  • Group A: Used the Compact Brain assistant.
  • Group B: Used the Mid-Size Brain assistant.
  • Group C: Used the Giant Brain assistant.

The participants could choose to ask the AI, read the PDF manual directly, or use a mix of both. They were paid extra if they got the answers right, so they were motivated to be accurate.

The Big Findings

1. The Power of the "Human-in-the-Loop"

The most important discovery was that humans working with AI were significantly better than AI working alone.

  • The Analogy: Imagine a GPS (the AI) that gives you directions. If you just follow the GPS blindly, you might miss a road closure. But if you have a human driver who can look at the map, question the GPS, and say, "Wait, that doesn't look right," the team gets to the destination much faster and more accurately.
  • The Result: Whether the AI was the "Compact," "Mid-Size," or "Giant" brain, the human-AI team crushed the scores of the AI working by itself.

2. Does a Bigger Brain Matter? (The Accuracy Test)

  • When the AI works alone: The "Giant Brain" (70B) was the best at finding answers on its own. The "Compact Brain" (3B) struggled more.
  • When the Human helps: This is where it gets interesting. When humans teamed up with the "Compact Brain," the humans helped fill in the gaps. The team's performance jumped so high that the gap between the small brain and the giant brain disappeared.
  • The Takeaway: A human partner can help a smaller, cheaper AI perform just as well as a massive, expensive one.

3. What Did the Humans Feel? (The Satisfaction Test)

The researchers asked the participants: "Did you like the assistant? Was it easy to use? Did you trust it?"

  • The Surprise: Even though the "Giant Brain" was technically the smartest when working alone, the humans didn't feel a huge difference between the three assistants when they were working together.
  • The Nuance: The only clear preference was that people felt the "Giant Brain" was slightly more reliable. However, for everything else (how easy it was to talk to, how much they liked it, how engaging it was), the "Compact," "Mid-Size," and "Giant" brains were rated almost the same.
  • The Analogy: It's like driving three different cars (a compact car, a sedan, and a luxury SUV) with a co-pilot. Even if the luxury SUV has a more powerful engine, once you have a co-pilot navigating, you might not feel like the luxury car is a better driving experience than the compact one.

Why This Matters

The paper argues that we often obsess over making AI models bigger and bigger to get higher scores on computer tests (benchmarks). But in the real world, where humans are actually using the tools:

  1. Collaboration is key: Humans + AI is a winning team, regardless of the AI's size.
  2. Bigger isn't always "felt" as better: Users didn't necessarily feel the massive, expensive models were much better to work with than the smaller ones.
  3. Efficiency: Since smaller models are cheaper and faster, and humans can help them perform just as well, we might not need to build giant, energy-hungry models for every job.

In short: You don't always need a supercomputer to solve a problem if you have a smart human to help you use a smaller, simpler tool. The human-AI partnership levels the playing field.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →