Shapley in Context: Explaining Financial Language with Domain Expertise
This paper investigates the use of Shapley values to explain large language models in financial applications, demonstrating through theoretical and empirical analysis that these attributions align with established domain knowledge and provide meaningful insights consistent with financial reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you've hired a super-smart, but mysterious, financial advisor (a Large Language Model or LLM) to read a company's report and tell you if it's a good investment or a risky gamble. The advisor gives you a score: "This is a 9 out of 10 on the risk scale!"
But here's the problem: The advisor won't tell you why. It just gives you the number. In the world of finance, where people lose real money and regulators demand answers, a "black box" that just gives numbers isn't good enough. You need to know: Was it the mention of "rising interest rates"? Was it the "debt"? Or was it just a generic warning about "employees"?
This paper is about building a fair, mathematical "magnifying glass" to figure out exactly which words in the report drove that risk score. The authors call this tool the Shapley Value.
The Core Idea: The Fair Pie Slicer
Think of the final risk score as a delicious pie. The ingredients (the words in the report) are what baked the pie. The Shapley Value is a method for slicing that pie so that every ingredient gets a fair share of the credit (or blame) based on how much it actually contributed.
If the word "Bankruptcy" was in the report, it probably took a huge slice of the pie. If the word "The" was there, it probably got a crumb. The Shapley Value calculates this fairly by testing the recipe over and over: "What happens to the pie if we remove 'Bankruptcy'?" "What if we remove 'The'?" "What if we have both?" By averaging all these scenarios, it assigns a precise value to every single word.
The "Financial Rulebook" (Axioms)
The authors didn't just want any fair slicer; they wanted one that respects the rules of the financial world. They created a set of "Financial Rulebooks" (which they call axioms) to test if their Shapley Value tool makes sense to a human expert.
Here are the rules they checked:
- The "Good Word" Rule (Monotonicity): If a word is clearly positive (like "Strong Earnings"), the tool must give it a positive score. If you swap "Strong" for "Weak," the score should drop. The tool passed this test.
- The "Diminishing Returns" Rule: Imagine you have one apple; you're happy. You get a second apple; you're happier, but not twice as happy. The tool understands that adding a second "Strong" word doesn't double the impact as much as the first one did. It passed this test too.
- The "Context Matters" Rule: The tool knows that "Stable Earnings" is great for a boring, safe company (like Coca-Cola), but might be less exciting for a tech company that needs to grow fast (like Amazon). The tool successfully adjusted the scores to reflect these different company personalities.
The Big Test: The 10-K Filing
To prove their tool works in the real world, the authors didn't just look at short sentences. They fed it massive, boring, legal documents called Form 10-Ks. These are the annual reports public companies must file with the government, containing a section called "Risk Factors."
They tested three very different companies:
- Silicon Valley Bank (SVB): Before it collapsed, their report was full of warnings about "Credit Risks" and "Liquidity." The tool looked at the report and said, "Hey, the 'Credit Risk' section is the biggest driver of the danger here." It correctly identified the specific words that predicted the bank's failure.
- Wells Fargo: This bank has a long history of regulatory scandals (like the fake accounts). The tool looked at their report and said, "The 'Regulatory Risks' section is the heavy hitter here." It correctly distinguished Wells Fargo's problems from SVB's problems.
- Crocs (The Shoe Company): They compared Crocs' report from 2020 (during the pandemic) to 2021. The tool noticed that the "Economy Risk" score went down (because the pandemic was easing) while the "Acquisition Risk" went up (because they bought a new shoe brand).
Why This Matters
The paper concludes that this "Fair Pie Slicer" (the Baseline Shapley Value) is a reliable way to open the black box of AI in finance.
- It's Consistent: If you run the test ten times, the ranking of the most important words stays the same.
- It's Accurate: If you remove the words the tool said were important, the risk score changes drastically. If you remove the unimportant words, the score barely moves.
In short, the authors built a mathematical method that doesn't just guess why an AI made a decision. It proves that the AI's decision-making logic aligns with how human financial experts actually think, making it safe to use these powerful AI tools for high-stakes money decisions.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.