How Much of a 10-K Matters? Aggregation-Dependent Value of Full-Text versus Risk-Factor Sentiment
This paper demonstrates that while full-text 10-K filings yield superior sentiment metrics for sector and portfolio-level predictions of returns and volatility, the narrower Item 1A risk-factor sections outperform at the individual firm level due to the interplay between document volume and available training signal, thereby establishing a supervised lexicon-learning approach as more effective than traditional dictionaries for regulatory disclosure analysis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the stock market as a giant, noisy ocean where the waves are prices and the currents are investor feelings. For decades, scientists trying to predict these waves have mostly listened to the chatter on the news—short, punchy stories about what companies are doing right now. But there's another, much deeper source of information that has been largely ignored: the massive, dry annual reports companies are legally required to file. Think of these reports as the company's official diary, a thick book filled with every detail of their past year and, crucially, a specific chapter where they must admit all the scary things that could go wrong in the future.
The big question this research tackles is: How do we turn these boring, legal documents into a "mood meter" for the market? Usually, computers just scan these texts with a pre-made list of "good" and "bad" words, like a teacher grading a paper with a fixed rubric. But this paper suggests that a smarter approach is to let the computer learn its own list of words by watching how the market actually reacts. The goal is to see if reading the whole annual report tells us more about future price changes and how shaky those prices might be (volatility) than just reading the specific section where companies list their fears.
The Great Report Card Experiment
In this study, the researchers acted like detectives trying to figure out which part of a company's annual report holds the most secrets about its future. They focused on 1,383 reports from 94 big technology companies (the kind of firms you might find in a tech-focused investment fund) spanning from 2006 to 2023.
They set up a clever game with two main rules:
- The Text: They compared reading the entire annual report (the "Full 10-K") against reading just one specific section called "Item 1A," which is dedicated entirely to listing risks and uncertainties.
- The Target: They trained their computer models to predict two different things: where the stock price will go next (returns) and how much the price will jitter or shake (volatility).
They tested this in three different ways: looking at the whole technology sector as one big group, looking at a small portfolio of the top 10 companies, and looking at a single company (Nvidia) on its own.
The Twist: Size Matters, But Only Sometimes
The results were a bit like a magic trick where the answer changes depending on how you look at it.
When looking at the big picture (the whole sector or a group of 10 companies):
Reading the entire annual report was the winner. The computer learned better from the massive volume of text in the full document. It was like trying to guess the weather by listening to a whole crowd of people; even if some people are talking about irrelevant things, the sheer number of voices helps you hear the real trend. The full report gave the computer more "training material" to figure out which words actually mattered.
When zooming in on a single company:
The magic flipped! Suddenly, reading just the risk section (Item 1A) became the better choice. Why? Because when you only have one company's history to learn from, the rest of the annual report is full of "boilerplate" stuff—standard legal jargon and financial numbers that don't change much from year to year. It's like trying to guess a single person's mood by reading their entire diary, including pages about what they ate for breakfast every day for ten years. That extra noise drowns out the important clues. But the risk section is short, punchy, and full of the specific worries that actually move the needle for that one company.
What the Words Actually Said
The researchers didn't just look at scores; they peeked at the words the computer decided were important.
- For the whole sector: The "good" words were about innovation and health (like "musk" or "torque"), while the "bad" words were about tech struggles.
- For the risk sections: These were surprisingly specific. The risk-only model spotted themes the full report missed, like "Supply Chain Risk" (words like "subcontractor" and "dependence") and even specific mentions of "Biotechnology" or "Cybersecurity." It seems the risk section is where companies whisper the specific dangers they are most afraid of, while the rest of the report is just shouting generalities.
The "Old Dictionary" Problem
The study also tested an old, popular method that uses a fixed list of "negative" words (like the Loughran-McDonald dictionary). The result was a bit funny: this old method was consistently wrong in a very predictable way. It was so negative that it actually moved in the opposite direction of the stock price. The researchers explain that this isn't because the dictionary is broken, but because it was designed to be super cautious. Since 10-K reports are legally required to be cautious, the dictionary sees "risk" everywhere, even when the market is feeling fine. This proves that a computer that learns from the data (the supervised approach used here) is much better than one that just uses a static list of words.
The Bottom Line
The paper concludes that there is no single "best" way to read a 10-K report. If you want to understand the mood of the whole tech industry, read the whole book. But if you want to understand the specific fears of one company, skip to the risk chapter. The key takeaway is that the amount of information you need depends entirely on how many companies you are looking at. The more companies you group together, the more text you need to find the signal; the fewer companies you look at, the more you need to cut out the noise and focus on the specific risks.
This research doesn't claim to have solved the stock market, but it provides a new, smarter toolkit for turning legal documents into useful signals, showing us exactly where to look for the truth depending on the size of the crowd we are trying to understand.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.