HGFX: a validated Python/JAX reproduction of the Hierarchical Gaussian Filter toolbox
HGFX is a validated Python/JAX toolbox that ensures numerical and behavioral parity with the reference MATLAB Hierarchical Gaussian Filter implementation while eliminating MATLAB dependencies, confirming GPU applicability, and transparently preserving known limitations.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
The human mind is constantly trying to make sense of a world that rarely stays the same. When we learn, we are not just absorbing facts; we are guessing how reliable those facts are and how quickly the rules of the game might change. This process of learning under uncertainty is a central puzzle in neuroscience. To study it, scientists use mathematical models that act like virtual laboratories, simulating how a brain might update its beliefs when faced with new information. One of the most trusted tools for this work is a specific computer program called the Hierarchical Gaussian Filter. For over a decade, this program has been the standard reference, written in a language called MATLAB, helping researchers understand everything from how we perceive sensory details to how we guess other people's intentions. However, relying on a single, older software environment can be a bottleneck, limiting who can use the tool and how easily it can be combined with modern, high-speed computing hardware.
A researcher has now created a new version of this essential tool, built from the ground up in Python, a language that is widely used in modern science and data analysis. Their goal was not to invent a new theory or a better way to learn, but to build a perfect digital twin of the original program. They wanted to prove that they could recreate the exact same scientific results without needing the original software. This was a delicate task. In complex calculations, tiny differences in how numbers are handled can sometimes grow into large errors, leading a computer to a different conclusion than another. The researcher treated this project as a rigorous test of faithfulness, checking every step to ensure their new code behaved exactly like the old one, right down to the smallest numerical details.
The researcher, led by Mohammad Ahmadkhanloo, released their new toolbox, named HGFX, as a direct replacement for the original. They did not simply translate the code; they treated the original program as a strict "oracle," a frozen standard against which every new calculation was measured. They tested the new system against two official demonstration workflows that are commonly used by scientists. In one of these tests, the original program is known to encounter a specific numerical difficulty where a calculation breaks down, while a slightly different version of the model handles it successfully. The new Python tool reproduced this exact behavior: it failed in the same way the original did when it was supposed to fail, and it succeeded when it was supposed to succeed. This was not a bug to be fixed, but a feature to be preserved, proving that the new tool understood the original scientific contract perfectly.
Beyond just copying the behavior, the researcher checked if the new tool could run on modern hardware. They tested it on physical graphics processing units, the powerful chips often used for artificial intelligence, and found that it produced results indistinguishable from the standard computer processor. The difference in the final calculation was so small it was measured in the range of one part in ten trillion, well within the limits of what is considered a perfect match. They also compared their work to another existing Python tool for the same problem. In the specific scenarios where a fair comparison was possible, the two tools agreed on the core predictions of the model, though they differed in how they handled certain edge cases involving extreme probabilities.
Crucially, the researcher was honest about what their new tool could not do. They ran a series of tests to see if the model could reliably recover hidden parameters from simulated data, a common requirement in scientific studies. In these tests, the new tool matched the original program's performance exactly, but that performance did not meet the high standards required to claim the model was fully successful at recovering those hidden values. Instead of hiding this limitation or claiming success, they explicitly reported it as a shared limitation of both tools. This approach ensures that future scientists using the new Python version will not accidentally believe they have solved a problem that remains unsolved.
The result is a tool that allows the scientific community to move away from a dependency on a single, older software environment without losing the reliability of decades of previous research. By providing a version that runs on modern systems and produces identical results to the gold standard, the researcher has opened the door for more scientists to use these powerful models. The work confirms that it is possible to rebuild complex scientific software in a new language without silently changing the science, provided the builder is willing to accept the original's flaws as part of the truth. This ensures that the next generation of discoveries in how we learn and perceive the world rests on a foundation that is both modern and rigorously faithful to the past.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.