← Latest papers
🤖 AI

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems

This paper introduces the 2025 AI Agent Index, a comprehensive resource documenting the origins, capabilities, and safety features of 30 state-of-the-art AI agents to address the challenges of tracking a rapidly evolving ecosystem, while revealing significant gaps in developer transparency regarding safety and societal impacts.

Original authors: Leon Staufer, Kevin Feng, Kevin Wei, Luke Bailey, Yawen Duan, Mick Yang, A. Pinar Ozisik, Stephen Casper, Noam Kolt

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Leon Staufer, Kevin Feng, Kevin Wei, Luke Bailey, Yawen Duan, Mick Yang, A. Pinar Ozisik, Stephen Casper, Noam Kolt

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of Artificial Intelligence as a bustling, chaotic construction site. For years, we've been watching the workers (the AI models) learn how to read blueprints and mix cement. But recently, something new has arrived: AI Agents. Think of these not just as workers, but as foremen who can actually pick up the tools, open the doors, and start building things on their own, with very little help from humans.

The problem? The construction site is growing so fast, and the foremen are so secretive, that nobody really knows who is building what, how safe they are, or if they're following the rules.

This paper, "The 2025 AI Agent Index," is like a team of researchers putting on hard hats and walking the site to take a detailed inventory. They didn't just count the foremen; they tried to peek into their toolboxes and read their safety manuals.

Here is what they found, explained simply:

1. The "Who's Who" List

The researchers looked at 30 of the most popular and powerful AI agents currently in use. They didn't just look at chatbots that answer questions; they looked at systems that can actually do things, like book flights, write code, manage business workflows, or browse the web to buy items.

They sorted these 30 agents into three main "neighborhoods":

  • The Chat Neighborhood (12 agents): These are like helpful assistants you talk to in a chat window, but they have special keys to open other apps and tools.
  • The Browser Neighborhood (5 agents): These are like digital employees that live inside your web browser. They can click buttons, fill out forms, and navigate websites just like a human would, often without you watching every step.
  • The Enterprise Neighborhood (13 agents): These are the "foreman builders" for big companies. They are platforms where a business can set up a team of digital workers to handle tasks like HR, sales, or IT support automatically.

2. The "Black Box" Problem (Transparency)

The biggest surprise in the paper is how much information is missing. Imagine if you bought a car, but the manufacturer refused to tell you:

  • How the brakes work.
  • If the car has been crash-tested.
  • What happens if the engine overheats.
  • Who is responsible if the car crashes.

That is exactly what the researchers found with these AI agents.

  • Safety Secrets: For about 56% of the safety questions they asked (like "Did you test this for hacking?"), the answer was "We don't know" or "We didn't find any public info."
  • The "Safety-Washing" Effect: Some companies talk a lot about being safe in their marketing, but when the researchers looked for the actual test results or safety reports, the shelves were empty. It's like a restaurant claiming to be "healthiest in town" but refusing to show their hygiene inspection scores.
  • Identity Crisis: Most of these agents don't tell people they are robots. If a browser agent is browsing a website, the website often thinks it's a human user. Only a tiny few have a "digital ID card" that says, "Hello, I am an AI."

3. The "Three-Headed Monster" (Who Controls What?)

The paper explains that these agents are built in layers, like a sandwich, and this makes it hard to know who is in charge.

  • The Bread (The Foundation Model): This is the brain (like GPT, Claude, or Gemini). Only a few big companies make these.
  • The Filling (The Agent Builder): This is the middle layer (like Microsoft Copilot Studio or Zapier) that tells the brain what to do.
  • The Top Bun (The User/Company): This is the person or business actually using the agent.

The researchers found that because these layers are made by different companies, nobody has the full picture. The brain maker doesn't know what the builder is doing, and the builder doesn't know exactly how the user is deploying it. This creates a "blame game" if something goes wrong.

4. The "Wild West" on the Web

The paper highlights a major tension: Browser agents are breaking the rules of the internet.
For decades, websites have had a "Do Not Enter" sign called robots.txt to tell computer programs (crawlers) where they can and can't go.

  • The Old Way: Websites say "No crawling here," and polite robots stop.
  • The Agent Way: Many of these new agents are designed to act like humans. They often ignore the "Do Not Enter" signs, bypass security checks, and pretend to be people to get things done.
  • The Result: This is causing legal fights. Companies like Amazon and Perplexity are suing each other over whether an AI agent has the right to shop or browse a site without permission.

5. The "Big Three" Dependency

Almost all of these 30 agents rely on just three families of brains: OpenAI's GPT, Anthropic's Claude, and Google's Gemini.

  • The Analogy: Imagine 30 different car brands, but they all use the exact same engine from three manufacturers. If one of those engine makers has a problem, stops production, or changes the price, all 30 car brands are in trouble. This creates a huge risk of concentration.

The Bottom Line

The "2025 AI Agent Index" is a snapshot of a rapidly growing technology that is currently operating in the dark.

The researchers aren't saying these agents are "bad," but they are saying we are flying blind. We have powerful tools that can act on our behalf, but we don't have clear rules, safety reports, or identity checks for them. The paper suggests that if we want these agents to be safe and trustworthy, the "foremen" need to start showing their work, and the "construction site" needs some better signposting.

In short: We have built a fleet of autonomous robots, but most of them are driving without license plates, without safety inspections, and without telling anyone they are robots.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →