← Latest papers
⚛️ phenomenology

LLM-Based FORM Code Generation with Verification-Driven Fine-Tuning

This paper addresses the lack of AI support for the FORM symbolic manipulation language by demonstrating that frontier large language models fail at zero-shot FORM tasks, then introduces a verification-driven fine-tuning pipeline using the FORM binary as an oracle to train a compact 8B-parameter model that outperforms much larger frontier models in code execution and correctness while preserving general reasoning capabilities.

Original authors: Bakar Chargeishvili

Published 2026-09-22
📖 4 min read🧠 Deep dive

Original authors: Bakar Chargeishvili

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the subatomic world, where particles collide and decay in fractions of a second, theoretical physicists rely on a specialized language to describe the mathematics of their discoveries. This language, known as Form, is not a general-purpose programming tool but a highly specific system designed to handle algebraic expressions of staggering complexity. When scientists calculate the behavior of particles using quantum field theory, the resulting formulas can contain millions or even billions of terms. General-purpose computer programs often crash when faced with such massive calculations, but Form is built to manage them by storing intermediate steps on a hard drive rather than in the computer's memory. For decades, this tool has been indispensable for uncovering the secrets of the universe, yet it remains difficult to learn. Its syntax is unique, its rules are unforgiving, and until now, there has been no artificial intelligence capable of helping a human write code for it.

This gap in technology presented a unique challenge for researchers at the Karlsruhe Institute of Technology. They asked a fundamental question: could a modern artificial intelligence, trained on vast amounts of internet data, learn to write code for a language that barely exists in that data? The answer they found was a definitive no. When they tested the most powerful, state-of-the-art AI models available—some containing hundreds of billions of parameters—these giants of artificial intelligence failed completely. Without a manual or a guide to look at, they could not generate a single line of valid Form code. To the AI, this language was effectively invisible, a zero-resource domain where the sheer size of the model mattered less than the absence of specific training.

To solve this, the researchers took a different approach. Instead of relying on a massive model to guess the rules, they built a specialized, smaller AI and taught it using a rigorous, self-correcting method. They created a pipeline where a powerful AI would generate candidate programs based on simple English instructions, but these programs were not trusted immediately. Instead, they were fed into the actual Form software to see if they ran correctly. If the code produced an error or the wrong result, it was discarded. Only the programs that passed every check—syntax, execution, and logical consistency—were kept. This process generated a verified library of over 4,600 examples, ranging from simple tutorials to complex physics calculations.

Using this high-quality, verified data, the team fine-tuned a compact AI model with eight billion parameters. The result was a specialist that outperformed the largest, most expensive models in the world. On a set of 664 specific calculation tasks, this small, specialized model solved 97.7% of them correctly. In contrast, the largest frontier models, even when given a syntax guide to read, managed to generate syntactically valid code that ran without error for only about 75% of the same tasks. The advantage was even more pronounced on open-ended tasks where the goal was described but the method was not specified; the specialized model successfully generated runnable code 83% of the time, while the best of the giant models managed this only 65% of the time.

Perhaps most surprisingly, the researchers found that this specialized training did not make the AI "forget" how to do other things. The model retained its general reasoning and coding abilities, showing only a tiny drop in performance on standard tests for math and logic. This suggests that teaching an AI a new, niche skill does not require sacrificing its general intelligence. The study demonstrates that for highly specialized scientific languages, the key to success is not the sheer size of the model, but the quality and verification of the data it learns from. By using the software itself as a teacher to verify every lesson, the researchers created a tool that can now assist physicists in writing the complex code needed to explore the fundamental laws of nature, lowering the barrier to entry for a language that has long been a bottleneck in high-energy physics.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →