A learned fungal ITS embedding does not outperform correctly configured alignment in open-world evaluation
This study demonstrates that a purpose-trained fungal ITS embedding does not outperform a correctly configured alignment-based approach in open-world evaluation, revealing that methodological flaws such as heuristic defaults, missing coverage filters, and data leakage—not the representation itself—were responsible for previously reported advantages of learned embeddings.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the hidden world beneath our feet, fungi are everywhere, yet they remain largely invisible to the naked eye. To understand this vast kingdom of life, scientists rely on a specific stretch of genetic code known as the ITS region. Think of this code as a unique barcode printed on every fungal species. When researchers collect a sample from soil or air, they sequence this barcode and try to match it against a massive library of known fungi to identify what they have found. However, nature is full of surprises. Often, the fungus in the sample is so new or so distant from known relatives that it does not match any entry in the library. The challenge for scientists is twofold: they must decide when a fungus is truly new and should be left unnamed, and they must place it as accurately as possible within the family tree of life, even if the exact species is unknown.
For years, the standard way to solve this matching problem has been to line up the new genetic sequence against every known sequence in the library, looking for the closest fit. This method, called alignment, is like checking every page of a dictionary to find the word that looks most like the one you are trying to spell. Recently, a newer, more high-tech approach has emerged. Instead of checking every page, scientists have begun using artificial intelligence to translate the genetic code into a mathematical point in a vast, multi-dimensional space. In this new system, similar fungi are placed close together, and different ones are pushed apart. The hope was that this learned system could spot new fungi and place them in the right family faster and more accurately than the old, slow method of checking every page.
A team of researchers set out to test whether this new artificial intelligence method truly outperformed the traditional approach. They built a custom AI model trained specifically to understand fungal genetics, using a massive, up-to-date collection of fungal sequences. To ensure a fair test, they designed a strict experiment where the AI and the traditional method were judged on the exact same set of unknown samples. They also took extraordinary care to seal their test data before the experiment began, ensuring that no part of the test could accidentally influence the training of the AI. This "firewall" meant that when the results were finally revealed, they would be a true measure of how well the methods worked on completely new data.
The researchers discovered that the outcome of the test depended entirely on how the rules were set. When they ran the traditional alignment method with its default, standard settings, the new AI model appeared to perform better. However, the researchers realized that the standard settings were not actually set up to do the best possible job. The traditional method had been skipping some potential matches and stopping its search too early. When the team reconfigured the traditional method to search exhaustively and carefully, checking every possibility without shortcuts, the results changed completely. The old method, now running at its full potential, became more accurate than the AI at identifying known fungi and placing new ones into the correct family groups.
The study revealed that the AI model had a hidden weakness: its performance changed depending on how many samples were fed to it at once. When the AI processed samples in large groups, the results were slightly different than when it processed them one by one. This inconsistency meant the AI was not truly stable. In contrast, the traditional method gave the same result every time, regardless of how the data was grouped. Furthermore, the researchers found that the AI's ability to spot new fungi was not actually better than the traditional method when measured at the specific points where a scientist would make a real-world decision. The AI did not find more new fungi, nor did it make fewer mistakes. In fact, the traditional method found more new fungi while making fewer errors.
Perhaps most surprisingly, the researchers found that the AI's success in earlier tests was partly due to a flaw in how the test was set up. Some of the "new" fungi in the test had actually been seen by the AI during its training, even though they were supposed to be hidden. When the researchers fixed this by removing those specific fungi from the training data and retraining the AI, the gap between the two methods widened. The traditional method pulled further ahead, showing a clear and statistically significant advantage. The AI's ability to compare different parts of the genetic code, a feature that was supposed to be its superpower, turned out to be no better than what the traditional method could do simply by looking at the shared parts of the code directly.
The final conclusion was clear: a carefully configured traditional method matched or exceeded the performance of the new artificial intelligence model in every category that mattered. The AI did not outperform the old method; it was the way the old method was being tested that had made it look weaker. The study showed that the choice of evaluation rules, such as how strictly the search was conducted and how the data was grouped, decided the winner. The researchers emphasized that these checks are simple and inexpensive to apply to any new method. They argued that before declaring a new artificial intelligence system superior, scientists must ensure that the traditional baseline is running at its best, that the data is truly new to the system, and that the results are stable under different conditions. In the end, the most reliable tool for identifying the unseen fungal world remained the one that had been used for decades, provided it was used correctly.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.