FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
The paper introduces FormalTCS, an expert-validated benchmark using frontier TCS research from 2025–2026 to reveal that current large language models struggle significantly with end-to-end automated research, primarily due to severe limitations in autoformalization and the ability to generate novel, high-quality mathematical claims.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Theoretical computer science is the study of how information is processed and the fundamental limits of what can be computed. It is the bedrock of modern technology, providing the mathematical rules that determine which problems can be solved by machines and how efficiently they can be solved. For decades, this field has relied on human mathematicians to craft precise definitions, construct logical arguments, and prove that their conclusions are unshakeable. Recently, a new kind of tool has emerged: large language models. These are artificial intelligence systems trained on vast amounts of text that can read, write, and reason about complex topics. As these systems have grown more powerful, researchers have begun to ask if they can do more than just answer questions or summarize text. Could they actually conduct scientific research on their own, discovering new truths about computation and proving them with the same rigor as a human expert?
A team of researchers from the Harbin Institute of Technology set out to answer this question by building a new testing ground called FORMALTCS. They wanted to move beyond simple quizzes or textbook exercises and see if these artificial intelligences could handle the messy, difficult reality of cutting-edge research. To do this, they gathered 175 real research problems from the most prestigious conferences in the field, covering topics like algorithms, complexity, and machine learning theory. These were not made-up problems; they were the actual core discoveries from papers accepted in 2025 and 2026. The researchers then worked with human experts to translate these discoveries into a format that a computer could verify, creating a rigorous standard against which to measure the artificial intelligence.
The results of this experiment revealed a stark divide between what these models can understand and what they can actually do. When the researchers asked the models to read a high-level summary of a research finding and explain the logic behind it in plain language, the models performed quite well. They could grasp the big picture and describe the strategy of a proof with reasonable accuracy. However, the moment the task shifted to translating those ideas into a formal, mathematical language that a computer could check, the models hit a wall. This process, known as autoformalization, is the step where a human researcher turns a vague idea into a precise set of rules and definitions. In this critical stage, the best artificial intelligence models managed to succeed only about 11.5 percent of the time. By contrast, when the researchers gave the models the precise mathematical definitions to work with, the models were able to construct the final proof about 28.6 percent of the time.
This gap suggests that the primary obstacle for artificial intelligence in this field is not the ability to follow a logical path once it is laid out, but the ability to build the path itself. The models struggle to take a natural language description of a problem and turn it into the specific, unambiguous definitions required for a computer to verify the solution. It is as if the models can read a map and understand the destination, but they cannot draw the map itself. Furthermore, when the researchers built a system that allowed the models to try to generate their own new research ideas from scratch, the results were even more telling. Out of 64 new claims generated by the system, only six were deemed novel and valuable enough by human experts to be worth pursuing. The rest were either too similar to existing work, lacked real significance, or simply could not be proven.
The study concludes that while these artificial intelligence systems are becoming better at understanding and discussing theoretical computer science, they are still far from being able to conduct independent research. They lack the "taste" or intuition required to identify which new ideas are truly worth exploring, and they struggle with the precise translation of ideas into formal rules. The path to fully autonomous scientific discovery in this field requires solving two distinct problems: improving the ability to translate ideas into rigorous mathematical language, and developing a deeper sense of what makes a research idea valuable. Until these hurdles are cleared, the role of artificial intelligence in this field will remain that of a powerful assistant rather than an independent researcher.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.