Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations
Forage V2 introduces an autonomous agent organization architecture that overcomes "denominator blindness" in open-world tasks by institutionalizing knowledge accumulation and transfer across model capabilities, thereby enabling weaker agents to achieve stronger performance, reduced costs, and faster convergence through shared, model-agnostic organizational memory.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are sending a team of explorers into a vast, unmapped jungle to find every single type of rare orchid that exists.
In the old way of doing things (what the paper calls Forage V1), you send one explorer. They find 50 orchids, get tired, and say, "I found them all! 100% complete!" But in reality, there were 200 orchids. The explorer didn't know how many there should be, so they just assumed their small pile was the whole forest. This is called "Denominator Blindness"—they can count what they found (the numerator), but they are blind to the total size of the search (the denominator).
Forage V2 changes the game. Instead of a lone explorer, it creates a Learning Organization. Here is how it works, using simple analogies:
1. The Two-Headed Team (The Evaluator and The Planner)
In V2, every mission has two distinct roles that never see each other's secret notes:
- The Planner (The Scout): Their job is to go out, find the orchids, and bring them back. They write the map and the collection strategy.
- The Evaluator (The Auditor): Their job is to stand back and ask, "Is this really everything? How many orchids should be out there?" They don't know how the Scout found them; they only judge the result.
Why separate them?
If the Scout judges their own work, they might think, "I tried really hard, so I must be done." By separating them, the Auditor stays honest. They can't be tricked by the Scout's effort; they only care about the truth.
2. The "Post-Mortem" (Learning from Mistakes)
In the old days, when a team returned from the jungle, they packed up, and all their hard-learned lessons were lost. The next team had to start from zero, making the same mistakes (like trying to climb a cliff that was too steep).
Forage V2 introduces a Post-Mortem. After every mission, the team sits down and writes a "Field Journal."
- Example: "Don't try to scrape data from TechPowerUp; they block robots."
- Example: "Wikipedia tables are messy; use a curated list instead."
These journals are saved in a Shared Library (the Knowledge Base).
3. The "Strong" vs. "Weak" Explorer (Knowledge Transfer)
This is the magic trick of V2.
- The Strong Explorer (Opus): A very smart, expensive AI model. It goes on 6 missions, learns everything, and builds a massive library of 54 "Field Journals."
- The Weak Explorer (Sonnet): A cheaper, less powerful AI model.
In the past, if you sent the Weak Explorer, they would struggle, take twice as long, and cost more money because they had to rediscover everything the Strong Explorer already knew.
With V2: You give the Weak Explorer the Strong Explorer's Field Journals before they leave.
- The Weak Explorer doesn't need to be as smart. They just need to read the handbook.
- They skip the mistakes. They don't waste time trying to climb the blocked cliff.
- The Result: The Weak Explorer, armed with the Strong Explorer's experience, performs almost as well as the Strong Explorer, but at half the cost and in half the time.
4. The "Institution" vs. The "Individual"
The paper argues that we shouldn't just try to make the AI "smarter" (like training a dog to be a genius). Instead, we should build a better organization (like a well-run company).
- Old Way: "Let's train one dog to be a super-genius."
- V2 Way: "Let's give every dog, even the average ones, a perfect rulebook and a strict manager who checks their work."
The system is designed so that even if the AI is "average," the structure (the rules, the separation of duties, the shared library) makes the result reliable.
5. The "Common Law" System
The knowledge in the library isn't a strict law that forces the AI to do exactly one thing. It's more like Common Law (legal precedents).
- The AI reads the past cases: "In 2023, we tried X and it failed."
- The AI decides: "Okay, I won't do X. But maybe I can try Y."
- This allows the AI to be flexible while still avoiding known pitfalls.
Summary: What Did They Prove?
- Knowledge Accumulates: Over 6 runs, the team's "Field Journal" grew from 0 to 54 entries. The team got smarter about what they were looking for, not just how to find it.
- Knowledge Transfers: A weaker AI, given the stronger AI's notes, stopped making the same mistakes. It found the same number of "orchids" (products) but spent half the money and time.
- Calibration: The weaker AI didn't just find more items; it actually started to estimate the total number correctly (e.g., "There are exactly 266 products"), whereas without the notes, it guessed wildly (e.g., "There are 400!").
The Big Picture:
Forage V2 shows that in a world where we don't know the "right answer" beforehand (like exploring the real world), we don't need super-intelligent robots. We need smart systems that remember their past, separate the doers from the judges, and pass their lessons down to the next generation. It turns a chaotic expedition into a reliable, learning organization.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.