Practical Principles for AI Cost and Compute Accounting
This paper proposes seven principles for designing AI cost and compute accounting standards to close technical loopholes, prevent strategic gaming, and ensure consistent regulatory implementation without discouraging responsible risk mitigation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
As artificial intelligence systems grow more powerful, governments are searching for a reliable way to decide which machines need strict supervision and which do not. The challenge is that these systems are complex, and their inner workings are often hidden inside the private servers of technology companies. To solve this, policymakers have turned to two measurable facts: how much money it costs to build a model and how much computing power it requires to learn. These numbers act as a proxy, a stand-in for the actual intelligence and potential danger of the system. The logic is straightforward: if a project requires a massive investment of resources, it likely possesses advanced capabilities that could pose risks to society. Laws in Europe and the United States are beginning to use these thresholds to trigger regulations, assuming that any project crossing a certain line of spending or computing power must be watched closely. However, a new paper by researchers Stephen Casper, Luke Bailey, and Tim Schreier argues that without clear rules on how to count these resources, the system is full of holes that clever developers could exploit.
The researchers set out to fix the technical ambiguities that currently allow companies to hide the true scale of their work. They identified that current proposals often fail to account for the entire journey of a model's creation. For instance, a developer might train a smaller, cheaper model by teaching it to mimic the answers of a much larger, more expensive "teacher" model. If regulations only count the final step, the massive cost of training the original teacher is ignored, making the new project look deceptively cheap. Similarly, companies might try to split their work across different legal entities or use open-source tools to bypass limits. To address these issues, the authors propose seven practical principles for designing accounting standards that are hard to game and easy to enforce.
The first and most critical principle is to count every dollar and every unit of computing power spent from the very beginning of a project up to the final system. This means including the work done to prepare data, the training of teacher models used for distillation, and even the failed experiments that were discarded along the way. The authors argue that limiting the count to only the final training run creates a "distillation loophole," where the true cost of a model's intelligence is masked. By requiring a full accounting of the upstream work, regulators can see the real investment behind a system's capabilities.
At the same time, the paper suggests that not everything should be counted. Developers should be allowed to exclude the costs of resources that are already freely available to everyone, such as open-source datasets or pre-trained models that were released publicly long before the current project began. This ensures that the rules focus on the proprietary effort a company is making, rather than penalizing them for using the shared foundation of the field. However, the authors add a crucial caveat: if a company releases a partially finished model to the public and then quickly picks it up to finish it, that initial work should still count. This prevents a strategy where a company tries to "reset" its accounting by briefly releasing a model before continuing its development in secret.
The framework also recognizes that some work is done specifically to make AI safer, not smarter. Activities like filtering out harmful content from training data or testing a model to ensure it refuses to help with criminal plans should be exempt from the cost calculations. The goal is to ensure that companies are not financially punished for doing the right thing and reducing risks to society. To make sure these exemptions are legitimate, the authors insist that companies must provide detailed, itemized reports. These reports would act like a financial audit, forcing developers to explain exactly what they did, why they did it, and how they calculated their numbers, including any estimates used when precise data was unavailable.
Finally, the paper advises that regulations should use two separate triggers: one for the total cost and another for the total computing power. Because these two metrics can sometimes be manipulated independently—such as by using cheap but computationally heavy data versus expensive but efficient human data—having both as safety nets makes it much harder to circumvent the system. The authors also emphasize that these standards cannot be static. As technology evolves and becomes more efficient, the thresholds and rules must be updated regularly, perhaps every few months, to remain effective. Without this adaptability, a rule that works today might become obsolete tomorrow, allowing dangerous systems to slip through the cracks simply because the accounting rules did not keep pace with the innovation.
The researchers conclude that while no accounting system can be perfect, these principles provide a solid foundation for creating rules that are transparent, consistent, and aligned with the public interest. By closing the loopholes that allow for strategic gaming and by encouraging responsible risk management, these guidelines could help governments oversee the most advanced artificial intelligence without stifling the smaller developers who do not pose the same level of risk. The ultimate aim is to create a regulatory environment where the true scale of an AI project is visible, ensuring that the most powerful systems receive the attention they require.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.