← Latest papers
📈 economics

Can we create a `race to the top' for weather forecasts to inform smallholder farmer decisions?

The paper proposes establishing standardized principles and protocols for evaluating agriculturally relevant AI weather forecasts to prevent a "race to the bottom" in quality and ensure that high-quality, tailored predictions effectively reach smallholder farmers in low- and middle-income countries.

Original authors: Colin Aitken, Michael K. Tippett, Pedram Hassanzadeh, Katherine Kowal, Rendani Mbuvha, John H. Marsham, Shruti Nath, Ousmane Ndiaye, Douglas J. Parker, Caroline M Wainwright, Michael Kremer, William R
Published 2026-10-02
📖 7 min read🧠 Deep dive

Original authors: Colin Aitken, Michael K. Tippett, Pedram Hassanzadeh, Katherine Kowal, Rendani Mbuvha, John H. Marsham, Shruti Nath, Ousmane Ndiaye, Douglas J. Parker, Caroline M Wainwright, Michael Kremer, William R. Boos

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

For millions of small-scale farmers in low- and middle-income countries, the difference between a bountiful harvest and a ruined season often comes down to a single decision: when to plant, when to water, or when to harvest. These choices rely heavily on knowing what the weather will do next. For decades, the promise of better weather forecasts has been held back by the immense cost and complexity of generating them. Traditional weather prediction requires massive supercomputers and teams of experts, making it difficult to produce tailored forecasts for specific regions or individual farms. However, a new wave of artificial intelligence has changed the landscape. These AI models can now produce high-quality weather predictions using a fraction of the computational power, opening the door for startups, researchers, and local agencies to create forecasts specifically designed for agricultural needs.

This technological shift brings a new challenge. Because the barrier to entry has lowered, the market is suddenly flooded with forecast products of varying quality. While some of these new models are genuinely accurate, others are flawed, misleading, or simply too unreliable to trust with a farmer's livelihood. The danger is that without a way to distinguish the good from the bad, farmers and the organizations that fund them may end up relying on cheap, inaccurate forecasts. If a farmer makes a costly mistake based on a bad forecast, they may lose faith in weather information entirely, causing the entire system to collapse. The core question is no longer just about making better forecasts, but about creating a system where quality can be proven and trusted.

A team of researchers from universities and institutions across the globe has proposed a solution to this problem: a set of clear rules and standards to evaluate weather forecasts before they reach the farmer. Their work, published in a recent paper, argues that the current lack of standards risks a "race to the bottom," where low-quality forecasts drive out the high-quality ones. To prevent this, they suggest a "race to the top," where forecasters compete to prove their accuracy through rigorous, transparent testing. The researchers do not claim to have invented a new forecasting model; instead, they have built a framework for verification. They outline a checklist of principles that any forecast intended for agricultural use must pass to be considered reliable.

The first and most critical principle is transparency. The researchers argue that a claim of accuracy is meaningless unless the method used to prove it is open for anyone to see. If a company says their forecast is better than the average, they must show exactly how they measured that improvement. This includes sharing the code and data used for the test, so independent experts can check for errors. The paper highlights a common pitfall where a model might look good on paper but fails in practice because the testing method was flawed. For instance, a model might be tested against a very weak baseline, making it look superior when it is actually worse than standard, freely available forecasts. The proposed rules require that new forecasts be compared against the best existing options, including open-source models and historical weather patterns, to ensure they offer a genuine improvement.

Another major focus is the concept of "out-of-sample" skill. In simple terms, this means testing a model on weather data it has never seen before. The researchers warn that many AI models are trained on vast amounts of historical data, and if they are tested on the same data they learned from, they may simply be memorizing the past rather than predicting the future. This is a form of overfitting. To avoid this, the proposed standards demand that forecasters clearly separate their training data from their testing data. They must also be honest about any decisions made during the testing phase, such as tweaking the model to perform better on a specific set of years. If a model is adjusted based on the test results, those results are no longer a fair measure of its ability to predict the unknown.

The paper also emphasizes that a forecast must be relevant to the actual decisions a farmer needs to make. A model might be excellent at predicting the total amount of rain over a month, but if a farmer needs to know if it will rain on a specific day to harvest their crops, that skill is useless. The researchers point out that many forecasts fail because they are not tested against the specific questions farmers are asking. They propose that forecasters must demonstrate how their product helps a user make a specific decision, such as choosing when to plant seeds. Furthermore, the language used in the forecast must be understood by the farmer. A statistic like "90% chance of rain" can be misinterpreted as a guarantee of a heavy storm, leading a farmer to take unnecessary precautions. The proposed standards suggest that forecasters test their messages with real users to ensure the information is clear and actionable.

The researchers illustrate the stakes with two cautionary tales. In one scenario, a government distributes a forecast from a leading global model because it scores well on standard metrics. However, the model has a hidden flaw: it frequently predicts light rain on days that are actually dry. While this doesn't ruin the overall score, it causes farmers to miss their narrow windows for harvesting, leading to crop loss. In another example, a group claims to predict floods by showing they got the dates right for the last three major floods. An agency invests millions in this system, only to discover the group had ignored all the times they predicted a flood that never happened. The model had predicted "ten of the last three floods," a classic case of cherry-picking data to hide a high failure rate. These examples show why a simple claim of success is not enough; the full picture of performance must be visible.

To address these issues, the authors propose a certification process. They envision a group of researchers, meteorological agencies, and non-profits agreeing on a set of principles that forecasters must follow. This group would train new teams on these standards, certify that the principles were followed, and build a consensus on how to update the rules as science evolves. The goal is to create a trusted seal of approval that funders and governments can rely on. If a forecast carries this certification, it means it has been tested against a rigorous checklist: the methods are transparent, the comparison is fair, the data is truly new, and the message is useful to the farmer.

The paper acknowledges that this is a starting point, not a final solution. The science of weather prediction is evolving rapidly, especially with the rise of artificial intelligence, and the rules will need to be updated as new challenges arise. However, the authors believe that establishing these baseline standards is essential to prevent the market from being flooded with unreliable products. By creating a system where quality is verified and transparent, they hope to build trust between forecasters and the farmers who depend on them. If successful, this approach could allow high-quality forecasts to reach hundreds of millions of people, helping them make better decisions and securing their livelihoods against the unpredictability of the weather. The path forward is not just about better technology, but about better trust.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →