← Latest papers
💬 NLP

MARCA: A Checklist-Based Benchmark for Multilingual Web Search

The paper introduces MARCA, a bilingual (English and Portuguese) benchmark featuring 52 manually curated questions with checklist-style rubrics to evaluate and compare the web-search capabilities, orchestration strategies, and multilingual performance of 14 large language models.

Original authors: Thales Sales Almeida, Giovana Kerche Bonás, Ramon Pires, Celio Larcher, Hugo Abonizio, Marcos Piau, Roseval Malaquias Junior, Rodrigo Nogueira, Thiago Laitz

Published 2026-04-17
📖 5 min read🧠 Deep dive

Original authors: Thales Sales Almeida, Giovana Kerche Bonás, Ramon Pires, Celio Larcher, Hugo Abonizio, Marcos Piau, Roseval Malaquias Junior, Rodrigo Nogueira, Thiago Laitz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of super-smart robots (Large Language Models, or LLMs) that can answer any question you ask. Usually, you just ask them, and they pull answers from their memory. But what if you need them to find brand new information on the internet, like "Who won the 82nd Golden Globes?" or "List all the current governors of Brazil"?

The problem is that these robots sometimes get lazy, make things up (hallucinate), or miss important details. Also, most tests only check if they are good at finding things in English. But what about Portuguese?

This paper introduces MARCA, a new "report card" designed to test how well these robots can search the web, read the results, and put together a perfect answer in both English and Portuguese.

Here is a simple breakdown of how they did it and what they found, using some everyday analogies:

1. The Test: The "Checklist" Game

Instead of just asking the robot, "Is the answer right?", the researchers gave the robots a very specific Checklist.

  • The Analogy: Imagine you hire a personal assistant to plan a dinner party. You don't just say, "Make it good." You give them a checklist:
    1. Buy 5 steaks.
    2. Get a bottle of red wine.
    3. Invite 4 friends.
    4. Turn on the oven.
  • The MARCA Twist: The researchers created 52 complex questions (like "List all the winners of the Golden Globes film categories"). For each question, they wrote a checklist of exactly what must be in the answer.
  • The Goal: If the robot forgets even one winner on the list, it gets a lower score. This stops the robot from giving a "vague but mostly correct" answer. It forces them to be precise.

2. The Two Ways to Work: "The Solo Artist" vs. "The Conductor"

The researchers tested the robots in two different ways to see how they handle big tasks.

  • Mode A: The Solo Artist (Basic Framework)
    • How it works: One robot tries to do everything alone. It searches, reads websites, and writes the answer all by itself.
    • The Metaphor: It's like one person trying to juggle 10 balls at once. They might drop a few.
  • Mode B: The Conductor (Orchestrator Framework)
    • How it works: One "Conductor" robot breaks the big question into smaller pieces and hires other "Sub-agent" robots to do the specific work. The Conductor doesn't search the web; it just manages the team.
    • The Metaphor: It's like a movie director. The director doesn't act, paint, or sew the costumes. They tell the actors, the painters, and the seamstresses what to do, then put it all together.
    • The Result: For most robots, being a "Conductor" worked much better. Breaking a big, messy task into small, focused jobs helped them get more items on the checklist correct.

3. The Language Gap: English vs. Portuguese

The researchers wanted to see if the robots were equally good at searching in English and Portuguese.

  • The Finding: It's a mixed bag.
    • Some robots are like polyglots: They perform just as well in Portuguese as they do in English.
    • Other robots are like tourists who only speak English: When asked to search in Portuguese, they struggle. They might try to translate the question into English in their head, search for English results, and then translate them back. This works fine for general topics, but fails when the answer only exists in Portuguese (like local Brazilian government details).
  • The Takeaway: Just because a robot is smart in English doesn't mean it's smart in Portuguese. The "internet" looks different in different languages, and the robots need to know how to navigate those specific streets.

4. The Price Tag: Better Answers Cost More

Finally, they looked at the cost.

  • The Analogy: Think of it like ordering food.
    • Cheap Meal: A quick, simple answer (Low cost, low accuracy).
    • Gourmet Meal: A detailed, perfectly researched answer with a full checklist (High cost, high accuracy).
  • The Result: Generally, if you want the robot to get a near-perfect score (90%+ on the checklist), you have to pay more. However, some robots are "efficient chefs"—they can make a great meal without costing a fortune, while others are expensive but still make mistakes.

Summary

MARCA is a new tool that stops us from just trusting robots blindly. It forces them to prove they did their homework by checking off specific items on a list.

  • Key Lesson 1: Giving robots a team (Orchestrator) usually helps them get better answers.
  • Key Lesson 2: Being good at English doesn't guarantee being good at Portuguese.
  • Key Lesson 3: If you need a perfect answer, you usually have to pay for it, but some robots offer better value than others.

This research helps developers choose the right robot for the job, especially if they are building tools for Portuguese speakers who need reliable, fact-checked information from the web.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →