← Latest papers
🤖 AI

Agents' Last Exam

This paper introduces Agents' Last Exam (ALE), a living benchmark developed with over 250 industry experts to evaluate AI agents on long-horizon, economically valuable real-world tasks across 13 industry clusters, aiming to bridge the gap between current benchmark success and meaningful GDP-relevant deployment.

Original authors: Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, Weishu Zhang, Tyler Zeng, Ying Yan, Bo Liu, Hanson Wen, Mingyang Xu
Published 2026-06-05
📖 5 min read🧠 Deep dive

Original authors: Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, Weishu Zhang, Tyler Zeng, Ying Yan, Bo Liu, Hanson Wen, Mingyang Xu, Xiaoyuan Liu, Zimeng Chen, Weiyan Shi, Amanda Dsouza, Vincent Sunn Chen, Patrick Bryant, Carl Boettiger, Yamini Rangan, Bradley Rothenberg, Kyle Steinfeld, Arvind Rao, Tapio Schneider, Georgios Yannakakis, Laure Zanna, Kaan Ozbay, Ida Sim, Tarek Zohdi, George Em Karniadakis, Jack Gallant, Teresa Head-gordon, Yushan Li, Wenxi Deng, Tao Sun, Huiqi Wang, Zhun Wang, Justin Xu, Chris Yuhao Liu, Yafei Cheng, Rongwang Hu, Aras Bacho, Shengcao Cao, Zengyi Qin, Yixiong Chen, Hengduan Fan, Hao Liu, Lin Zeng, Shashank Muralidhar Bharadwaj, Litian Gong, Yingxuan Yang, Maojia Song, Ruheng Wang, Zongzheng Zhang, Honglin Bao, Shuo Lu, Jianhong Tu, Zhonghua Wang, Zheng Zhang, Zijiao Chen, yanqiong Jiang, Zhendong Li, Bohan Lyu, Chang Ma, Peiran Xu, Benran Zhang, Shangding Gu, Haoyue Hua, Haoyang Li, Wanzhe Liao, Chengzhi Liu, Junbo Peng, Haoran Sun, Zechen Xu, Bo Chen, Jiayi Cheng, Yi Jiang, Keying Kuang, Yuan Li, Youbang Pan, Ziyan Rao, Alexander Schubert, Yifan Shen, Vincent Siu, Xiatao Sun, Kangqi Zhang, Xiaopan Zhang, Yuchen Zhu, Ishaan Singh Chandok, Lei Ding, Jingxuan Fan, Andrew Glover, Jiaming Hu, Yiran Hu, Wenbo Huang, Zixin Jiang, Haoran Jin, Lukas Kim, Ming Liu, Yang Liu, Alireza Rafiei, Xuhuan Shen, Kunyang Sun, Sophia Sun, Ting Sun, Eric Wang, Yixin Wang, Hanwen Xing, Sihan Xu, Yuzheng Xu, Zhongxing Xu, Zhiling Yan, Boqin Yuan, Ruiqi Zhang, Yifan Zhang, Zibo Zhao, Liana, Santanu Bosu Antu, Haoyue Bai, Carlo Bosio, Joseph Cavanagh, Patricia Cavazos-Rehg, Tianxing Chen, Xuewen Chen, Yipu Chen, Zhu Chenyu, Chen Dai, Stefano De Castro, Yunfu Deng, Kaustubh Dhole, Jiayuan Ding, Chenchen Du, Zhehang Du, Hao Fan, Run-ze Fan, Hengyu Fu, Shi Gu, Yifan Gu, Charlie Guo, Baihe Huang, Baixiang Huang, Rimika Jaiswal, Zhihan Jiang, Ran Jin, Erin Kasson, Xin Lan, Joseph Lee, Deren Lei, Chenyu Li, Daofeng Li, Haitao Li, Hongwei Li, Jingyan Li, Xiao Li, Yi Li, Yinsheng Li, Yuangang Li, Zhixu Li, Wenyu Liang, Longtai Liao, Kevin Qinghong Lin, AndyZeyi Liu, Che Liu, Jiaming Liu, Kaiyuan Liu, Xuan Liu, Pan Lu, Wenbo Lv, Yicheng Lv, Qiuyang Mang, Kyle Montgomery, Yuzhou Nie, Ruoxi Ning, Jorin Overwiening, Xu Pan, Layna Paraboschi, Core Francisco Park, Justin Purnomo, Swati Rajwal, Scott Rankin, Bixuan Ren, Yiren Rong, HaoYang Shang, Ventus Shaw, Fiona Shen, Jiawei Shen, Minqi Shi, Qiu Shi, Huaxiu Yao, Tianneng Shi, Jonah So, Vladislav Susoy, Hannah Szlyk, Haocheng Wang, Jialu Wang, Wei Wang, Xinyu Wang, Zehao Wang, Dowling Wong, Angela Wu, Dehao Wu, Fangyu Wu, Mengyuan "Millie" Wu, Yu Wu, Yuchen Wu, Yuhao Wu, Qingpo Wuwu, Weihang Xiao, Yongyi Xiong, Fan Xu, Ruiling Xu, Mingxuan Yan, Benjamin Yang, Jirong Yang, Sen Yang, Xiaoli Yang, Yushi Yang, Haoran Ye, Xiaohu Yu, Zhengming Yu, Chenlong Zhang, Chi Zhang, Hanning Zhang, Hanwen Zhang, Junge Zhang, Kunpeng Zhang, Song Zhang, Wenjin Zhang, Wenshuo Zhang, Ying Zhang, Yizhi Zhang, Brian Zhao, Qijian Zhao, Yimin Zhao, Yuhaohua Zheng, Liwei Zhou, Tianyue Zhou, Sichen Zhu, Siqi Zhu, Yan Zhu, Yishu Zhu, Jierui Zuo, Chonghao Cai, Helena Casademunt, Wenjia Chen, Benjamin Cheng, Nawen Deng, Rao Fu, Tianfu Fu, Yifan Han, Ren He, Zhenyu He, Qiao Jin, Lang Lang, Yuetai Li, Sylvia Liu, Lu Lu, Qing Lu, Subhabrata Mukherjee, Yunqi Ouyang, Yin Ren, Dawei Shi, Haoran Wu, Zhiyue Wu, Hannah Yao, Zhuoran Yi, Jenny Yu, Rhea Zhan, Hang Zhou, Blake Zhu, Junfan Zhu, Alan Yuille, Yang Liu, Russell Alan Poldrack, Jiachen Li, Zhenglu Li, Molei Tao, Jing Huang, Wenqi Shi, Costas Spanos, Lichao Sun, Chenguang Wang, Orson Xu, Zhen Dong, Hector Gomez, Aylin Caliskan, Ali Emami, Haimin Hu, Zhi Li, Lihui Liu, Murphy Niu, Yi Shao, Jianxin Sun, Mikko Tolonen, Ting Wang, Sanjiv Das, Yanjun Gao, Wenbo Guo, Erika J Schneider, Zhiyong Lu, Mark Mueller, Radha Poovendran, Somayeh Sojoudi, Dawn Song

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've been training a robot chef for years. You've tested it on quizzes about recipes, asked it to name ingredients, and even had it follow simple instructions like "chop the onion." On these tests, the robot gets perfect scores. It seems like a culinary genius.

But then, you ask it to actually cook a full, complex dinner for a busy restaurant during a rush hour. Suddenly, the robot burns the sauce, forgets the order, or gets confused by the stove. It turns out, being good at talking about cooking isn't the same as being good at doing the work.

This is exactly the problem the paper "Agents' Last Exam" (ALE) is trying to solve.

The Problem: The "Quiz" Trap

For a long time, AI researchers have been testing their AI agents (smart computer programs) on benchmarks that are like multiple-choice quizzes. These tests are great at measuring what an AI knows, but they don't measure what an AI can do in the real world.

The authors argue that just because an AI can pass a math test or a coding quiz, it doesn't mean it can actually run a business, design a bridge, or manage a hospital schedule. The gap between "passing the test" and "making money" is huge.

The Solution: The "Last Exam"

To fix this, the team created ALE (Agents' Last Exam). Think of this not as a quiz, but as a final, real-world job interview for AI.

Instead of asking, "What is the capital of France?" or "Write a Python loop," ALE gives the AI a real, messy, multi-day project that a human professional would actually do.

  • The Tasks: The exam covers 55 different job fields (like engineering, medicine, finance, and video game design).
  • The Difficulty: The tasks are "long-horizon," meaning they take hours or days to complete. They require the AI to open different software programs, click buttons, write code, check files, and fix its own mistakes, just like a human employee.
  • The Source: These aren't fake tasks made up by researchers. They are real projects that 250+ industry experts actually did in their jobs. The experts submitted their past work, and the team turned them into tests.

How the Exam Works

Imagine the AI is sitting at a computer in a virtual office.

  1. The Assignment: The AI gets a real-world task, like "Design a mold for a car part" or "Analyze these medical X-rays."
  2. The Tools: The AI has to use real software (like 3D modeling tools, financial spreadsheets, or medical imaging software) just like a human would. It has to click menus, type commands, and manage files.
  3. The Grading: This is the clever part. The exam doesn't rely on a human teacher to look at the result and say, "Looks good!" That's too slow and subjective. Instead, the exam uses automated checkers.
    • If the task was to build a 3D model, the computer checks if the dimensions are exactly right.
    • If the task was to write a financial report, the computer checks if the numbers add up correctly.
    • If the AI fails a "gate" (like making a mistake that would break a machine), it gets a zero immediately, no matter how pretty the rest of the work looks.

The Results: The AI is Still in School

The paper tested the world's smartest AI agents on this exam. The results were sobering:

  • The "Easy" Tier: Even the best AI agents only passed about 30% of the tasks that are considered "near-term" (easier ones).
  • The "Hard" Tier: On the most difficult tasks (the "Last Exam" tier), the pass rate was 0%. The AI couldn't do them at all.
  • The Gap: An AI that gets 82% on a standard coding test (Terminal-Bench) only gets about 25% on this real-world exam.

The authors call this the "Last Exam" because it represents the final hurdle. If an AI can pass this, it means it's ready to actually do the job and replace a human worker. Right now, the paper says, the AI is still far from passing.

Why It Matters

The paper isn't just about making a harder test; it's about changing the goal.

  • Current Goal: Make AI score 100% on quizzes.
  • New Goal: Make AI score 100% on real jobs that generate economic value (GDP).

The authors believe that by focusing on these real-world, verifiable tasks, we will stop building AI that is just good at talking and start building AI that is good at working. Until the AI can pass this "Last Exam," it's not ready for the real workforce.

Summary Analogy

Think of the current AI benchmarks as driving theory tests. You can memorize all the rules of the road and get a perfect score. But ALE is the driving test where you have to actually drive a car through heavy traffic, parallel park, and navigate a construction zone without hitting anything.

The paper says: "Our AI drivers are great at the theory test, but they are still crashing in the real driving test. Let's stop celebrating the theory scores and start teaching them how to drive."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →