← Latest papers
🧬 biology

Avoiding the explicit inverse of the pedigree relationship matrix among genotyped animals in approximate reliabilities of single-step genomic predictions

This study evaluates five computational strategies that successfully avoid the prohibitive memory and time costs of explicitly inverting the pedigree relationship matrix for single-step genomic predictions, with the Woodbury identity and stochastic Hutchinson approaches emerging as the most efficient methods for scaling approximate reliability calculations to large genotyped populations while maintaining high numerical accuracy.

Original authors: Gabriel Campos, Vinícius Silva Junqueira, Marcos Jun Iti Yokoo, Henry Carvalho, Fernando Flores Cardoso

Published 2026-08-20
📖 3 min read☕ Coffee break read

Original authors: Gabriel Campos, Vinícius Silva Junqueira, Marcos Jun Iti Yokoo, Henry Carvalho, Fernando Flores Cardoso

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

In the world of animal breeding, the goal is to improve herds by selecting the best parents for the next generation. To do this effectively, breeders rely on estimated breeding values, which are predictions of an animal's genetic potential. However, a prediction is only as good as its reliability, a measure of how precise that estimate is. For decades, calculating this precision was manageable when breeders worked only with family trees, or pedigrees. But the field has shifted toward single-step genomic predictions, a powerful method that combines family history with actual DNA data. This integration allows for much more accurate breeding decisions, but it introduces a massive computational hurdle. The DNA data links every genotyped animal to every other, creating a dense web of connections that makes the standard mathematical tools for calculating reliability too large to fit into a computer's memory. When a population grows beyond a few tens of thousands of genotyped animals, the traditional method of solving these equations simply crashes, leaving breeders without a way to know how much they can trust their best candidates.

A team of researchers set out to solve this bottleneck by finding new ways to calculate these reliability scores without ever building the massive, unwieldy matrix that causes the computer to fail. Working with data from the PROMEBO Angus breeding program in Brazil, which included nearly 434,000 animals with over 24,000 of them genotyped, the team tested five different strategies. Their goal was not to invent a new genetic theory, but to find a smarter way to do the math. They compared these new approaches against the traditional, exact method, which serves as the gold standard but is impossible to run on such large datasets. The researchers wanted to see if they could skip the step of explicitly creating the huge matrix, instead using mathematical shortcuts to get the same answer with far less memory and time.

The results showed that all five new strategies could reproduce the results of the gold-standard method with extreme accuracy. When the researchers compared the reliability scores generated by their new methods against the benchmark, the numbers matched almost perfectly, with correlations exceeding 99 percent. This means that the shortcuts did not sacrifice precision; they simply avoided the computational dead end. Among the five approaches, two stood out for their efficiency. One method, which uses a specific mathematical identity to rearrange the problem, and another that uses a statistical sampling technique, both reduced the memory required by about 80 percent. Instead of needing nearly 18 gigabytes of memory to run the calculation, these methods required only about 3.5 gigabytes. This reduction is critical because it means the calculation can be performed on standard servers rather than requiring specialized, expensive supercomputing hardware.

The speed of these methods varied, but one of the sampling-based approaches was particularly impressive. It completed the calculations in roughly the same amount of time as the traditional method, but without the memory crash. This is a significant finding because it suggests that as breeding programs grow to include hundreds of thousands or even millions of genotyped animals, the traditional method will become impossible to use, while these new strategies will remain practical. The researchers found that for smaller populations, the old method is still the fastest, but once the number of genotyped animals passes a certain threshold, the new methods become the only viable option. The study confirms that breeders can continue to use the most advanced genomic tools without hitting a wall of computational limits, ensuring that the precision of their genetic evaluations keeps pace with the size of their data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →