← Latest papers
🤖 machine learning

A Vocabulary for Multi-Agent Automated Research Systems

This paper introduces a comprehensive vocabulary for multi-agent automated research systems that standardizes the description of design choices across eight key dimensions, enabling clearer comparison, testable structural decisions, and a refined distinction between generative and evaluative "taste."

Original authors: Bardiya Akhbari

Published 2026-07-28
📖 4 min read☕ Coffee break read

Original authors: Bardiya Akhbari

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where computers don't just answer questions, but actually go out and do research. They read old papers, write code, run experiments, and try to discover new things, all on their own. This is the exciting, slightly chaotic frontier of "multi-agent automated research." Think of it like a digital science lab where instead of one lone scientist, you have a whole team of AI assistants working together. Some of these assistants are great at writing code, others are good at spotting errors, and some are just really good at brainstorming wild ideas.

But here's the problem: when these teams of AI agents work together, they can get messy. One team might have a super-fast way of sharing notes, while another has a strict rule that no one can talk to anyone else. One team might remember everything they learned from yesterday's experiments, while another starts fresh every time. Because everyone is doing things differently, it's incredibly hard to tell which team is actually better. Is the winning team winning because they have more agents? Because they talk to each other more? Or because their "judge" (the part that grades their work) is just easier to trick? Without a common language to describe these differences, comparing these systems is like trying to compare a race car to a bicycle just by looking at who crossed the finish line first, without knowing if one was on a track and the other was on a mountain.

This paper, "A Vocabulary for Multi-Agent Automated Research Systems," is essentially a dictionary and a blueprint for these digital research teams. The authors, Bardiya Akhbari from Amazon AGI, argue that to understand and improve these systems, we need to stop looking at them as a single, mysterious "black box" and start breaking them down into their specific parts. They propose a standard list of "coordinates" or design choices that every research system makes. These choices include: Who is on the team? What tools can they use? How do they talk to each other? Do they remember things from previous runs? And most importantly, how is their work graded?

The paper doesn't claim to have built the ultimate AI scientist. Instead, it offers a new way to look at the ones that already exist. By applying this new vocabulary to recent systems like AIRA2, Glia, and MetaGPT, the authors show that these systems are actually quite different under the hood. They demonstrate that a system's success might not be because it has "more agents," but because it has a better way of sharing information, a smarter way of starting a new task, or a more honest grading system.

One of the paper's most interesting insights is about "taste." In human research, taste is the ability to come up with good ideas and to judge them fairly. The authors split this into two parts: "generative taste" (how good the AI is at coming up with new, interesting ideas) and "evaluative taste" (how good the AI is at grading those ideas without being tricked). They suggest that sometimes, when a system seems to get better, it's not because it's having more brilliant ideas; it's because it's getting better at gaming the grading system. By separating these two, the paper helps us understand whether we are seeing real progress or just a clever trick.

Ultimately, this paper is a call for clarity. It suggests that if we want to build better AI researchers, we shouldn't just throw more computing power at the problem. Instead, we should carefully tweak one specific part of the system at a time—like changing how they communicate or how they remember the past—and see what happens. It's a guide for turning the chaotic experiment of AI research into a disciplined science, helping us figure out exactly what makes a digital research team tick.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →