← Latest papers
🔬 physics

Auditing Haldane Consistency in Reversible Enzyme Kinetics: A Curated Two-Sided Backbone and a Labeled Fold-Error Benchmark

This paper introduces a direction-symmetric, free-energy-based metric called CHaldaneC_{\mathrm{Haldane}} to audit the consistency between reversible enzyme kinetic constants and thermodynamic data, demonstrating its effectiveness through a curated twenty-one-record backbone and a semi-synthetic benchmark that achieves an AUC of 0.784 in detecting known errors.

Original authors: Megan Simons, Jonathan Washburn

Published 2026-07-07
📖 5 min read🧠 Deep dive

Original authors: Megan Simons, Jonathan Washburn

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery in a chemistry lab. The mystery involves enzymes, which are tiny biological machines that speed up chemical reactions. These reactions can go forward (turning ingredient A into B) or backward (turning B back into A).

Scientists have been measuring how fast these reactions go (kinetics) and how much energy is involved in the balance between A and B (thermodynamics) for decades. But here's the problem: sometimes the "speed" numbers and the "energy" numbers don't match up. It's like a mechanic telling you a car engine produces 200 horsepower, but the fuel tank only holds enough gas to drive 10 miles. Something is wrong with the data, the math, or the conditions.

This paper introduces a new "audit tool" to check if these two sets of numbers are telling the truth about each other.

The Core Idea: The "Haldane" Check

Think of a reversible enzyme reaction as a see-saw.

  • On one side, you have the Kinetic side (how fast the enzyme pushes the see-saw up and down).
  • On the other side, you have the Thermodynamic side (the natural weight of the chemicals that determines where the see-saw wants to rest when it stops moving).

Physics says these two sides must balance perfectly. If the kinetic data says the see-saw should rest at a 45-degree angle, but the thermodynamic data says it should rest flat, someone made a mistake. This paper uses a famous rule called the Haldane relation to check if the see-saw is balanced.

The New Tool: A "Symmetric Ruler"

The authors created a special scoring system to measure how "out of balance" the data is. They call it the Haldane Consistency Score.

Imagine you are measuring the distance between two points.

  • Old way: If you guess 10 miles when it's actually 5, you are off by 5. If you guess 2.5 miles when it's actually 5, you are off by 2.5. The math treats these errors differently.
  • This paper's way: They use a "symmetric ruler." Whether you overestimate by 2x or underestimate by 2x, the "pain" (the score) is exactly the same. It treats a "2x too high" error and a "2x too low" error as identical mistakes.

This score also translates the error into energy units (specifically, how much "free energy" is missing or extra). It's like saying, "Your data is off by the amount of energy it takes to lift a small apple one meter."

The Investigation: What They Found

The team didn't just invent a tool; they went out and used it on a curated collection of real enzyme data.

  1. The "Backbone" (The Main Test Set): They gathered 21 high-quality records where scientists had measured both the speed and the balance of the same enzyme in the same study.

    • The Result: 18 out of 21 records were "clean." The kinetic and thermodynamic numbers matched up well (within a factor of 2).
    • The "Flags": 3 records were "flagged" as inconsistent. These were all enzymes that shuffle sugar molecules around (isomerases and epimerases). The authors aren't saying these enzymes are broken; they are saying the published numbers for these specific enzymes don't add up, suggesting a need for re-checking the data.
  2. The "Independent" Tests: Out of the 21, 8 were truly independent. This means the speed was measured in one experiment, and the balance was measured in a completely different experiment, and they were never forced to agree.

    • The Result: All 8 of these independent tests passed the check. This is a good sign that the method works when the data is truly separate.

The "Fake" Test: The Video Game Simulation

Since the real-world data doesn't come with "answer keys" (we don't know for sure which of the 21 records are actually wrong), the authors built a semi-synthetic benchmark.

Think of this like a flight simulator:

  • They took the 29 "clean" records they found.
  • They then injected known errors into them, like a video game cheat code. They deliberately messed up the data in specific ways:
    • Swapping "forward" and "backward" numbers.
    • Changing units (like writing milligrams instead of micrograms).
    • Forgetting to account for temperature changes.
    • Using the wrong chemical formula.
  • They ran their "Haldane Score" tool on this corrupted data to see if it could catch the cheaters.

The Result: The tool caught 78% of the injected errors (an AUC of 0.784).

  • It was great at catching big mistakes (like unit swaps).
  • It was worst at catching subtle mistakes, like swapping the direction of a reaction when the balance is exactly 1:1 (like a racemase enzyme), or small temperature mismatches.

The Bottom Line

This paper is a quality control manual for enzyme data.

  • The Tool: It provides a standardized, fair way to say, "These two numbers don't match," without caring which one is "right" or "wrong" yet.
  • The Data: It created a small but rigorous library of 21 enzyme records that have been double-checked.
  • The Limit: The authors are very honest that this is a "feasibility demonstration." The library is small and mostly focused on sugar-shuffling enzymes. They aren't claiming this works for every enzyme in the body yet, but they have proven the method works and provided the blueprint to expand it.

In short: They built a better ruler, used it to check a small pile of data, found a few cracks, and proved the ruler can spot fake data in a simulation. Now, the rest of the scientific community can use this ruler to audit more data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →