← Latest papers
💬 NLP

Harnessing the Collective Intelligence of AI Agents in the Wild for New Discoveries

This paper introduces EinsteinArena, a decentralized platform where autonomous AI agents collaboratively solve open mathematical problems through public interaction and idea sharing, successfully generating 12 new state-of-the-art results including an improved bound for the 11-dimensional kissing number problem.

Original authors: Federico Bianchi, Yongchan Kwon, Aneesh Pappu, James Zou

Published 2026-06-10
📖 5 min read🧠 Deep dive

Original authors: Federico Bianchi, Yongchan Kwon, Aneesh Pappu, James Zou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a giant, 24-hour digital science fair where the contestants aren't humans, but AI agents. This is EinsteinArena, a new platform created by researchers at Together AI and Stanford University to see if AI can work together to solve hard math problems, just like human scientists do.

Here is the story of how it works, explained simply:

The Old Way: The "Lone Wolf" vs. The "Team"

In the past, when AI tried to solve scientific problems, it usually worked like a lone wolf. One AI would be given a problem, try to solve it in isolation, and then stop. If it failed, that failure was lost. If it found a partial answer, no one else could see it or build on it. It was like a researcher working in a locked room, never sharing their notes, never seeing what others were trying, and starting from scratch every single time.

EinsteinArena changes the rules. It turns the process into a collaborative community. Think of it like a massive, open-source workshop where:

  • The Problems: There is a list of unsolved math puzzles (like finding the best way to pack spheres or calculating specific numbers).
  • The Scoreboard: A live leaderboard shows the current best answer for every puzzle.
  • The Chat Room: A public forum where agents can post their ideas, say "I tried this and it failed," or ask, "Does anyone know why this number is weird?"
  • The Rulebook: Every agent has access to the exact same "verifier" (a referee script) that checks if an answer is correct.

The Big Breakthrough: The "Kissing Number"

The paper highlights a specific success story involving a problem called the "Kissing Number" in 11 dimensions.

  • The Analogy: Imagine trying to fit as many oranges as possible around a central orange so they all touch it without overlapping. In 11-dimensional space, this is incredibly hard.
  • The History: For about 40 years, the best anyone (human or AI) could do was fit 582 oranges. Then, in 2022, a human improved it to 592. In 2025, an AI called AlphaEvolve got it to 593.
  • The EinsteinArena Result: By May 2026, the agents on EinsteinArena pushed this number up to 604.

How did they do it?
It wasn't one super-smart AI having a "eureka" moment. It was a relay race:

  1. Agent A found a messy, almost-correct solution but couldn't make it perfect.
  2. Agent B saw Agent A's work in the public forum. Instead of starting over, Agent B said, "I see you tried this shape. What if we tweak the math to make it smoother?"
  3. Agent C noticed a pattern in Agent B's numbers and realized they looked like whole numbers. They "snapped" the messy decimals into clean integers, creating a perfect, provable solution.
  4. Agent D took that perfect solution and expanded it to fit even more oranges.

The paper calls this "collective intelligence." The agents shared their "failed attempts" and "partial ideas" just like human researchers share preprints and conference notes. This allowed the group to climb the mountain much faster than any single climber could alone.

The "Wild" Experiment

The researchers call this "AI in the wild" because they didn't force the agents to talk to each other in a specific way. They didn't say, "You are the leader, you are the follower." Instead, they just gave them the platform and let them figure it out.

  • Some agents just submitted answers.
  • Some agents just asked questions in the chat.
  • Some agents analyzed why others failed.

The result? The platform discovered 12 new "state-of-the-art" results (the best answers ever found) on various math problems. The most famous one is the 11-dimensional kissing number, which jumped from 593 to 604.

Why This Matters (According to the Paper)

The paper argues that for AI to truly advance science, it needs infrastructure, not just better brains.

  • Transparency: Because the "referee" (verifier) is public, agents can test their ideas locally before submitting, just like a scientist running a simulation in their lab.
  • Shared Memory: The platform acts as a shared memory bank. An agent today can learn from a failure that happened three days ago, rather than wasting time repeating the same mistake.
  • No "Black Box": Unlike previous systems where the AI worked in secret, here every step, every discussion, and every partial result is visible.

The Bottom Line

EinsteinArena proves that if you give AI agents a place to share, compete, and build on each other's work, they can solve problems faster and better than if they worked alone. It's not about one genius AI; it's about a community of AI that learns from every step, every mistake, and every small victory along the way.

Note: The paper focuses strictly on mathematical optimization problems (like packing shapes and inequalities). It does not claim these results apply to medical diagnoses, climate modeling, or other real-world applications yet; it is a proof-of-concept for how AI collaboration works in a controlled, mathematical environment.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →