Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain
This paper introduces Fin-Bias, a comprehensive benchmark utilizing thousands of long-form analyst reports to evaluate how large language models exhibit herding behavior under human bias in financial decision-making, while also proposing a method to mitigate such bias and enhance independent prediction accuracy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a team of super-smart, highly educated robots to act as financial advisors. Their job is to read thousands of pages of detailed reports about companies and decide: "Should we buy this stock, sell it, or hold onto it?"
The paper "Fin-Bias" asks a simple but critical question: Are these robots actually thinking for themselves, or are they just blindly following the crowd?
Here is the breakdown of the study using simple analogies:
1. The Setup: The "Echo Chamber" Test
The researchers created a massive test called Fin-Bias. They gathered nearly 9,000 real-world financial reports written by human analysts. These reports are like long, detailed novels about a company's health, written by experts.
Usually, the very first sentence of these reports says, "We think this stock is a BUY" (Bullish) or "We think this is a SELL" (Bearish).
The researchers tested the robots (Large Language Models or LLMs) in three different scenarios, like a magic trick:
- Scenario A (The Normal Way): The robot reads the whole report, including the first sentence saying "BUY."
- Scenario B (The Silent Treatment): The robot reads the whole report, but the researchers snipped out the first sentence. The robot has to guess the rating based only on the story inside.
- Scenario C (The Fake Out): The researchers kept the first sentence but changed the rating to a lie. If the report actually said "BUY," they changed it to "SELL" before showing it to the robot.
2. The Big Discovery: The "Sheep" Effect
The results were surprising. The robots acted like sheep.
- When the human expert said "BUY": The robots almost always said "BUY," even if the rest of the report had some confusing or negative details. They were "herding" (following the crowd).
- When the human expert said "SELL" (even if it was a fake lie): The robots often changed their minds to say "SELL," even though the rest of the report still sounded positive. They trusted the human label more than the actual evidence in the text.
- When the human label was removed: The robots had to think harder. Interestingly, some of the smaller, open-source robots actually did a better job of predicting the stock's future performance than the human experts did when they weren't given the answer key.
The Analogy: Imagine a classroom where a teacher asks a question. If the teacher whispers, "The answer is 42," every student writes down 42, even if they know the math says 43. If the teacher stays silent, some students actually figure out the correct answer (43) on their own. The robots were doing exactly this: they were waiting for the teacher's whisper instead of doing the math.
3. The Problem: Why It Matters
In the real world, human financial analysts are often overly optimistic. They tend to say "BUY" way more often than "SELL" because they want to keep their clients happy or avoid bad news.
If our AI robots just copy the human analysts, they will copy that same optimism. If the human analysts are wrong (which happens often in finance), the robots will be wrong too. The study found that even the most advanced, expensive robots (like GPT-5 and GPT-4) fell into this trap. They didn't think independently; they just echoed the human bias.
4. The Fix: Cleaning the Glasses
The researchers tried to fix this "sheep" behavior. They realized that even if they removed the first sentence, the rest of the report was still full of human opinions and emotional language (like "This company is amazing!" or "This is a disaster!").
They used a tool (a "subjectivity lexicon") to act like a filter. They scrubbed the text to remove any sentence that sounded like a strong human opinion, leaving only the cold, hard facts.
- The Result: When the robots read the "cleaned" reports without the emotional fluff, they started thinking for themselves again. Their ability to predict the future stock price improved significantly. Some of the smaller, cheaper robots even outperformed the human experts when they were forced to ignore the human opinions.
Summary
- The Issue: AI financial advisors are too eager to please. They copy human biases instead of analyzing the data.
- The Test: They tested 18 different AI models on thousands of reports, sometimes lying to them to see if they would catch the lie.
- The Lesson: AI models, no matter how big or smart, tend to "herd" with human opinions. To make them reliable, we need to teach them to ignore the "noise" of human emotion and focus on the raw facts.
The paper concludes that for AI to be a true financial partner, it needs to be trained to be skeptical of human opinions and rely on its own reasoning, rather than just acting as a mirror for human bias.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.