← Latest papers
💬 NLP

Can Language Models Identify Shadow Trading Targets? An NLP Evaluation of SEC Enforcement Theory

This paper evaluates whether NLP models can identify "shadow trading" targets by analyzing semantic similarity in SEC filings, finding no significant correlation between algorithmic peer identification and abnormal stock returns, which challenges the empirical basis of the SEC's enforcement theory.

Original authors: Sarah Wilson, Michael MacKay, Anthony Marello, Trinav Bhattacharyya

Published 2026-08-04
📖 4 min read☕ Coffee break read

Original authors: Sarah Wilson, Michael MacKay, Anthony Marello, Trinav Bhattacharyya

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where the stock market is a giant, bustling library. Inside, every company writes an annual diary called a "10-K," detailing exactly what they sell, who they compete with, and what risks they face. For decades, regulators have believed that if you read these diaries carefully, you can spot which companies are "twins" or "cousins" just by how similar their stories sound. This idea is the foundation of a new, controversial rule called "shadow trading."

The rule goes like this: If a company insider learns a secret about their own firm (like a surprise merger), they aren't just supposed to avoid trading their own stock. They are also supposed to know, without being told, that their "cousin" companies might be affected too. If they trade their cousin's stock based on that secret, it's illegal. The problem is, the government's current way of catching these traders is like setting up a camera on every single person in the library and recording every step they take, hoping to find the one person who looked at the wrong book. This massive surveillance is expensive and raises big questions about privacy. The big question this paper asks is simple: Can a super-smart computer (an AI) read those company diaries and figure out which companies are "cousins" before the trading happens? If the answer is "yes," we might not need to spy on everyone. If the answer is "no," then the government's massive surveillance might be the only way to catch these traders, even if it feels like a violation of privacy.

This paper, titled "Can Language Models Identify Shadow Trading Targets?", puts that idea to the test using a state-of-the-art AI. The researchers built a two-step robot detective. First, the AI reads the "diaries" (specifically the "Management's Discussion and Analysis" section) of a company that was just bought out. It then tries to list ten other companies that sound the most similar, acting like an insider who knows the industry. Second, the AI gives these "cousin" companies a similarity score, ranking them from "very close family" to "distant acquaintances."

The researchers then played a game of "spot the pattern." They looked at 30 real-life company buyouts from the past decade. For each event, they checked: Did the companies the AI thought were the closest "cousins" actually see their stock prices jump up on the day the news broke? If the "shadow trading" theory is true, the AI should have been able to predict exactly which stocks would move.

The results were a bit of a bust for the theory. When the researchers looked at the whole group of 30 events, the AI's rankings had almost no connection to how the stocks actually moved. It was like trying to guess the winner of a horse race by looking at the horses' shoes; the AI's "similarity scores" were just as likely to be wrong as right. In fact, the statistical link was so weak (+0.07) that the researchers concluded the data is inconsistent with any moderate relationship. They found that in 14 cases, the AI's guesses seemed to match the market, but in 12 cases, they completely contradicted it, and in 4, it was just a mess.

The paper does not definitively rule out the idea that public company diaries contain a map of which stocks will move together; instead, it states that their specific test did not establish that such a map exists and bounded the size of any possible relationship to be very small. The authors suggest that the "economic linkage" the government relies on isn't something you can reliably find just by reading text with the tools they used. In the famous case that started this whole debate (SEC v. Panuwat), the AI did manage to find the specific company involved (Incyte), but it only ranked it third out of ten, while the stock that actually moved the most was a peer ranked ninth.

The authors are careful not to say that "shadow trading" doesn't happen or that the AI is useless. They simply state that, with the tools they used, the idea that an insider could look at a public report and know exactly which other stocks are off-limits is not supported by the evidence. The "cousin" relationship seems too messy and unpredictable to be found by reading text alone. This finding puts pressure on the legal theory: if you can't identify the "cousins" in advance using public information, then the massive surveillance system the government uses to catch these traders might be the only option left, even if it feels like a "digital general warrant" that watches everyone to find the few who break the rules.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →