Classical baselines outperform released deep-learning ITS classifiers, which collapse on ITS2 where predictions follow the flanking regions
This study demonstrates that while deep-learning classifiers for fungal ITS sequences achieve high accuracy on full-length references, they fail dramatically on the commonly used ITS2 subregion by relying on flanking regions rather than the barcode itself, causing classical baseline methods to significantly outperform them in realistic environmental metabarcoding scenarios.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the microscopic world of fungi, scientists rely on a specific stretch of genetic code to identify species, much like a librarian uses a barcode to sort books. This code, known as the internal transcribed spacer or ITS, sits between other genes in the fungal DNA and varies enough between species to serve as a unique identifier. For decades, researchers have used this region to catalog the invisible fungal life that surrounds us, from the mushrooms in a forest to the molds growing in a home. However, modern technology has made it possible to sequence only a small, convenient fragment of this code—specifically a section called ITS2—rather than the entire genetic record. This shortcut is faster and cheaper, allowing scientists to process thousands of samples at once, but it raises a critical question: can the most advanced computer programs, trained on the full genetic record, still recognize the fungus when they only see this tiny fragment?
Recently, a team of researchers set out to test the limits of artificial intelligence in this field. They examined two powerful deep-learning classifiers, sophisticated computer models that had been trained on millions of fungal sequences to predict species names with high accuracy. These models were celebrated for their ability to identify fungi with over 90 percent accuracy, but they were trained on the complete, full-length genetic records. The researchers wanted to know if these same models would work when fed the short, partial sequences that environmental surveys actually produce. To find out, they created a fair test using thousands of fungal sequences, comparing the performance of the artificial intelligence against older, more traditional methods that rely on direct comparison rather than complex learning.
The results revealed a severe failure in the artificial intelligence. When the researchers fed the models the full-length genetic records, the deep-learning classifiers performed respectably, though they were still slightly less accurate than the traditional methods. However, the moment the input was restricted to just the short ITS2 fragment, the performance of the artificial intelligence collapsed. On this short fragment, the deep-learning models dropped to an accuracy of roughly 18 to 28 percent, a catastrophic failure compared to the traditional methods, which remained robust and accurate at nearly 90 percent. The artificial intelligence models, which had been touted as the future of fungal identification, essentially stopped working when faced with the very data they were most likely to encounter in the real world.
The researchers then investigated why this happened, suspecting that the models were not actually reading the barcode they were supposed to identify. They conducted a clever experiment where they took the short ITS2 fragment from one fungus and surrounded it with the flanking genetic regions from a completely different type of fungus. If the models were truly reading the ITS2 barcode, they should have identified the original fungus. Instead, the models overwhelmingly identified the donor fungus, the one whose surrounding genetic material had been attached. In one case, the model correctly identified the donor's family for nearly 64 percent of the queries, while identifying the actual fungus in less than 1 percent of cases. This proved that the models were not looking at the barcode itself but were instead relying on the surrounding genetic context, which was missing entirely from the short environmental samples.
This discovery explains why the models failed so dramatically. The artificial intelligence had learned to recognize fungi by looking at the entire genetic record, including the regions before and after the barcode. When presented with the short fragment alone, the models were effectively blind, unable to find the features they had been trained to recognize. Worse still, the models did not admit their confusion. Their confidence scores lost the ability to distinguish between familiar fungi and novel ones, meaning the output provided no warning that the prediction was likely wrong. In the case of one model, it remained overconfident, giving researchers a false sense of security despite being mostly incorrect. The study concludes that the reported high accuracy of these deep-learning tools is an illusion created by testing them on data that does not match the real-world conditions of environmental surveys. For the tools to be useful, they must be retrained or redesigned to understand that the short fragments they receive are fundamentally different from the full records they were built on, or else scientists must rely on the simpler, more reliable methods that have stood the test of time.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.