← Latest papers
🤖 AI

ASI-Bench: At the Dawn of Artificial Superintelligence

The paper introduces ASI-Bench, a comprehensive benchmark comprising 60 project-level research tasks across 11 scientific domains designed to evaluate AI's ability to conduct autonomous, innovative scientific research by progressively reducing human guidance, revealing that current state-of-the-art systems remain heavily dependent on human input and are far from achieving true artificial superintelligence.

Original authors: Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, C
Published 2026-08-19
📖 5 min read🧠 Deep dive

Original authors: Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The journey toward a machine that can think like a human has long focused on how well computers can memorize facts, solve known puzzles, or follow a set of instructions to reach a specific answer. For years, the most advanced artificial intelligence systems have excelled at these tasks, absorbing vast libraries of human knowledge and applying them with incredible speed. But there is a different kind of intelligence that remains largely out of reach: the ability to step into the unknown, to look at a messy, unexplained problem, and to figure out not just the answer, but the very path to finding it. This is the realm of scientific discovery, where the goal is not to recall a formula but to invent one, to design an experiment, and to turn a vague question into a verified result. The question driving a new study is whether current machines can truly do this on their own, or if they are still entirely dependent on humans to hold their hands through every step of the process.

A team of researchers from universities and institutes around the world has built a new testing ground called ASI-Bench to answer this question. They created sixty complex research projects across eleven different fields of science, ranging from physics and chemistry to biology and engineering. Each project was designed to mimic a real scientific investigation, complete with raw data, a specific goal, and a requirement to produce a working solution. The brilliance of the design lies in how they tested the machines. For every single research project, they ran the same AI system four times, but they changed the amount of help they gave it. In the first scenario, the AI was handed a complete, step-by-step manual on exactly how to solve the problem. In the second, they were told only the general method to use, like being told to "use a hammer" without being shown how to swing it. In the third, the AI was given only the problem and the data, with no instructions on how to proceed, forcing it to decide the method itself. Finally, in the fourth scenario, they added a layer of confusing, irrelevant information to see if the machine could stay focused on the real task.

The results reveal a stark reality about the current state of artificial intelligence. When the machines were given full instructions, they performed reasonably well, scoring an average of about fifty-one out of a possible hundred. However, as soon as the researchers removed the step-by-step guidance and asked the AI to figure out the procedure on its own, the performance plummeted. When the AI had to choose the method itself, the average score dropped to roughly twenty-seven. This sharp decline suggests that while today's systems are excellent at following a path that humans have already paved, they struggle immensely to build the road themselves. The biggest hurdle was not choosing the right tool, but rather translating a general idea into a working, detailed plan. Even when the AI knew which method to use, it often failed to construct the necessary steps to make it work, causing scores to drop significantly when the detailed instructions were removed.

The researchers also discovered that adding extra, distracting information did not confuse the machines nearly as much as removing the instructions did. When the AI was given the same difficult task but with a pile of unrelated facts and chatter added to the prompt, its performance remained almost exactly the same as when it had no instructions at all. This indicates that the primary weakness of current systems is not a lack of focus or an inability to filter noise, but a fundamental inability to operate without a human-defined roadmap. The study involved eighteen different combinations of AI models and software agents, and even the most advanced configuration, which used the most powerful reasoning available, managed to score only about fifty-two when left to its own devices. This is a modest success, but it falls far short of the reliability required for a machine to conduct independent scientific research.

The cost of this independence is also high. When the AI had to figure out the steps on its own, it consumed significantly more computing power and time than when it was simply following orders. In some cases, the system used nearly sixty percent more computing resources just to attempt the same task without a guide, often getting stuck in loops or trying inefficient paths. This suggests that true autonomy is not just a matter of intelligence, but of efficiency; without human guidance, the machines waste vast amounts of effort trying to solve problems they could have solved quickly if they had been told how. The study concludes that we are still far from the dawn of artificial superintelligence, a state where machines can explore the unknown and create new knowledge without human intervention. Instead, we are in a phase where machines are powerful assistants that can execute complex tasks, but they still rely heavily on humans to define the strategy.

This new benchmark is not intended to be a final judgment, but rather a starting point for the scientific community. The researchers have made their sixty tasks and their testing methods available to the world, inviting other scientists and engineers to contribute new problems and to test their own systems against these challenges. They believe that by continuously adding new, difficult research projects and refining the tests, the community can track progress more accurately. The goal is to move beyond simple tests of memory and logic to a future where we can measure whether a machine can truly think like a scientist, capable of navigating the unknown and turning a blank page into a discovery. For now, the data shows that while our machines are getting smarter, they are still waiting for us to tell them what to do next.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →