Disagreement Is Not Debiasing: Human–AI Views in Portfolio Decisions
This paper demonstrates that while adding an independent AI forecast to human analyst predictions improves probabilistic accuracy by refining uncertainty estimates through disagreement, it fails to correct shared systematic biases or enhance target-based portfolio decisions without calibration against realized outcomes.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the high-stakes world of investing, making a decision often feels like trying to hit a moving target while standing on a shifting platform. Investors rely on forecasts to guess how much money a company will make, but these guesses are rarely perfect. For decades, a standard approach has been to combine the opinions of human experts with the calculations of computer models, hoping that the mistakes of one will cancel out the mistakes of the other. This idea, known as forecast combination, suggests that if you take two imperfect predictions and average them, the result should be more accurate than either one alone. However, this logic assumes that the two sources are making different kinds of errors. If both the human and the machine are looking at the same documents and are influenced by the same optimistic mood, they might make the exact same mistake at the same time. In such a case, simply averaging their views does not fix the problem; it just reinforces the error.
This reality forms the backdrop for a new study conducted by researchers at Xiamen University Tan Kah Kee College, who set out to test whether an artificial intelligence system could correct the systematic biases of human financial analysts. The researchers focused on the Chinese stock market, a vast and active arena where thousands of companies are tracked by professional analysts. They gathered decades of data, looking at the earnings forecasts made by human experts and comparing them against forecasts generated by a large language model, a type of advanced AI that reads the same company reports but is strictly forbidden from seeing the analysts' own predictions. The goal was to see if the AI could act as an independent second opinion that cleans up the human error, or if the two sources were simply biased in the same direction, rendering their disagreement useless for fixing the core problem.
The study began by acknowledging a well-known fact: human financial analysts are consistently too optimistic. In the data examined, these experts predicted earnings that were higher than what actually happened in about three-quarters of the cases. The researchers then introduced the AI, which was fed the same raw financial data but was designed to be informationally independent, meaning it could not peek at the human consensus. A natural assumption might be that if the AI and the human disagree, the difference reveals the human's bias. If the human says a company will earn a lot and the AI says less, one might assume the human is wrong. However, the researchers found that this intuition is flawed. Because both the human and the AI were reading the same underlying reports, they shared a common blind spot. They were both biased in the same direction, just by different amounts.
The core discovery of the paper is that disagreement between two biased sources does not automatically correct the bias. When two forecasters share a common error, their disagreement only tells you how much they differ from each other, not where the truth actually lies. To find the truth, you need to look at what actually happened in the past, not just at the difference between the two current guesses. The researchers built a sophisticated model to test this. They first adjusted both the human and the AI forecasts based on historical data, essentially teaching the model to recognize and subtract the typical over-optimism found in each source. Once this calibration was done, they combined the two views to see if the result was better than using just the human forecast alone.
The results were nuanced and challenged the hope that adding an AI simply improves everything. At the level of pure prediction accuracy, the combination of the calibrated human and AI forecasts did improve the statistical description of uncertainty. The model became better at describing the range of possible outcomes, largely because the disagreement between the two sources helped the model understand when it should be more cautious. When the human and the AI disagreed, the model correctly interpreted this as a sign that the situation was uncertain and widened its safety margins. This improved the probabilistic accuracy of the forecast, making it a better tool for understanding risk.
However, when the researchers tested whether this improved forecast led to better investment decisions, the story changed. They simulated a portfolio strategy that required a specific, absolute target for returns. In this scenario, adding the AI forecast provided no detectable benefit over using the calibrated human forecast alone. The reason lies in the nature of the shared bias. While the AI helped the model understand the spread of possible outcomes, it could not fix the level of the forecast. Because both the human and the AI were still slightly too optimistic on average, the combined forecast remained too high. For a decision that depends on hitting a specific number, this small remaining error was enough to prevent any improvement. The disagreement between the two sources could not substitute for the hard work of calibrating against real-world results.
The study also revealed that the AI forecast itself was not a neutral, unbiased oracle. Like the human analysts, the AI was also biased upward, predicting earnings that were higher than what actually occurred. This finding is crucial because it dismantles the idea that an AI is automatically a "truth-teller" that can correct human error. Instead, the AI was another source of information that carried its own systematic flaws. The researchers found that the AI's bias was about one percentage point, distinct from the analyst bias of 1.45 percentage points. This means that simply adding an AI to the mix does not magically remove the optimism that plagues financial forecasting.
Ultimately, the paper demonstrates that the value of a second opinion depends entirely on the type of decision being made. If the goal is to rank companies from best to worst, the shared bias does not matter, because both the human and the AI are likely to rank them in the same order. But if the goal is to hit a specific target, like a minimum return rate, that shared bias becomes a critical obstacle. The research shows that you cannot rely on the difference between two forecasts to fix a shared error; you must learn from the actual outcomes of the past. The study concludes that while artificial intelligence can improve the understanding of uncertainty and the description of risk, it cannot replace the need for rigorous calibration against historical reality. In the end, a collaborative intelligence design must separate the task of learning from outcomes from the task of gathering new information, recognizing that disagreement is a signal of uncertainty, not a cure for bias.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.