What Counts as Compute? An Empirical Analysis of Accounting Conventions in Compute-Optimal Transformer Training
This paper demonstrates that while varying accounting conventions for counting parameters and compute significantly alter the fitted exponents of the compute-optimal training frontier, the resulting practical impact on training efficiency is modest due to the flatness of the loss landscape, with the specific accounting choices being predictable from simple architectural ratios.