← Latest papers
🤖 AI

Characterizing initial human-AI proof formalization workflows

This paper investigates the initial impact of AI on human proof formalization workflows through mixed-methods research, revealing that while users desire high-level human control, access to AI tools significantly improves formalization accuracy and encourages flexible, collaborative engagement between mathematicians and AI.

Original authors: Katherine M. Collins, Simon Frieder, Jonas Bayer, Jacob Loader, Jeck Lim, Peiyang Song, Fabian Zaiser, Lexin Zhou, Shanda Li, Sam Looi, Joshua B. Tenenbaum, Umang Bhatt, Adrian Weller, Jose Hernandez-
Published 2026-06-04
📖 4 min read☕ Coffee break read

Original authors: Katherine M. Collins, Simon Frieder, Jonas Bayer, Jacob Loader, Jeck Lim, Peiyang Song, Fabian Zaiser, Lexin Zhou, Shanda Li, Sam Looi, Joshua B. Tenenbaum, Umang Bhatt, Adrian Weller, Jose Hernandez-Orallo, Cameron E. Freer, Valerie Chen, Ilia Sucholutsky

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine mathematics as a massive, intricate construction project. For centuries, human mathematicians have been the architects and builders, drawing up blueprints (proofs) to prove their ideas are sound. But there's a catch: checking if a blueprint is perfectly error-free is incredibly hard, even for the experts. Sometimes, tiny mistakes slip through, leading to shaky foundations.

Enter AI, the new high-tech tool that promises to help build and check these blueprints faster. But before we just hand the keys over to the robots, the researchers in this paper asked a crucial question: "How do human builders actually want to use these tools, and how do they work with them in real life?"

Here is a breakdown of their findings, using simple analogies:

1. The Survey: What Builders Want vs. What They Fear

The researchers first asked a group of mathematicians (from students to professionals) what they thought about AI.

  • The Desire for Control: Most people didn't want AI to take over the whole construction site. They didn't want a robot architect designing the building from scratch while they watched. Instead, they wanted to remain the foreman. They wanted the AI to handle the heavy lifting, the repetitive brick-laying, and the tedious paperwork, but they wanted to keep the "big picture" decisions and the creative direction in their own hands.
  • The Trust Issue: Even though people loved the idea of AI saving time, they were frustrated by its reliability. It's like having a very enthusiastic intern who is great at finding tools but sometimes hands you a hammer when you asked for a screwdriver. Many users said, "I use it, but I have to double-check everything because I can't trust it 100% yet."
  • The "Swiss Army Knife" Approach: People didn't want just one AI tool. They wanted a toolbox. If one tool was good at finding existing blueprints but bad at writing new ones, they'd switch to a different tool for that specific job.

2. The Experiment: Putting Theory to the Test

To see how this actually works, the researchers ran a controlled experiment. They gave seven mathematicians a set of math problems (some easy, some very hard) and asked them to translate these problems into a strict, computer-readable language called Lean.

  • The Setup: The participants had to do this twice.
    • Week 1: They had to do it alone, with no AI help (just a text editor).
    • Week 2: They could use any AI tools they wanted (chatbots, code assistants, search engines).
  • The Result: When the mathematicians were allowed to use AI, they made fewer mistakes. The AI helped them get the "blueprint" right more often.
  • The Time Factor: Interestingly, using AI didn't necessarily make them finish faster. It was a trade-off. They spent time figuring out which tool to use and checking the AI's work, but the final result was more accurate.

3. How They Actually Worked: The "Hybrid" Dance

The most interesting part was watching how they used the tools. They didn't just let the AI do the work and walk away. They developed a hybrid workflow:

  • The "Human-in-the-Loop" Style: Most participants acted like a conductor leading an orchestra. They would ask the AI for a suggestion (like a specific theorem or a code snippet), listen to the AI, and then decide: "Yes, that fits," or "No, that's wrong, try again."
  • The "Tool Switcher": One participant, for example, used one tool to find a specific rule, another tool to write the repetitive parts of the code, and a third tool to rewrite a section to make it look better. They were constantly juggling different tools to get the job done right.
  • The "AI-Heavy" Style: A few participants let the AI do almost all the heavy lifting, only stepping in to fix the final product. This worked, but it required the human to have a very strong ability to spot errors at the end.

The Bottom Line

The paper concludes that the future of math isn't about humans vs. AI, or even humans replacing AI. It's about partnership.

Think of it like driving a car with a very advanced GPS. The GPS (AI) can tell you the fastest route and warn you about traffic, but the human (the mathematician) still needs to hold the steering wheel, decide when to take a detour, and make sure the car doesn't drive off a cliff.

The study shows that when humans keep the "steering wheel" (strategic control) but let the AI handle the "navigation" (finding patterns and checking details), they build better, more accurate proofs. The key isn't just how smart the AI is; it's how well humans learn to adapt and use these new tools to fit their own style of working.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →