Dynamic Topic Modeling for Cross-Corpus Temporal Analysis
This paper proposes a Dynamic Embedded Topic Model framework that establishes a shared topic space across multiple corpora with corpus-specific residual adaptations, enabling stable and interpretable cross-corpus temporal comparisons while preserving unique lexical variations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
For decades, social scientists and historians have relied on computers to make sense of vast libraries of text, from old newspapers to business magazines. The goal is to find the hidden themes that run through these documents, much like a librarian sorting books by subject rather than by author. A common tool for this is called a topic model, which scans thousands of pages to group words that frequently appear together, revealing what people were talking about at a given time. When researchers want to see how these conversations change over years or even centuries, they use a dynamic version of this tool that tracks how the meaning of these word groups shifts as time passes. However, a major hurdle has always existed when trying to compare different collections of text. If you analyze a collection of business magazines and a separate collection of labor union reports, the computer usually treats them as two completely different worlds. It might find a topic about "money" in the business magazines and a topic about "wages" in the union reports, but it has no built-in way to know that these are actually two sides of the same coin. Traditionally, researchers had to run the analysis separately and then try to manually match the results afterward, a process that often failed to line up the themes correctly, leaving the comparison shaky and unreliable.
A team of researchers at Columbia University has developed a new way to solve this problem, allowing computers to compare different types of text collections while keeping the themes perfectly aligned across time. Instead of analyzing each collection separately and trying to force a match later, their method builds a single, shared map of ideas first. Imagine this shared map as a common coordinate system, like a standard set of latitude and longitude lines, that covers all the different text collections at once. The computer learns the main themes and how they evolve over a ninety-seven-year period using this combined map. Once this shared foundation is established, the researchers allow each specific collection—whether it is business news, labor reports, or general historical writing—to make small, controlled adjustments to the map. This lets the business magazines use their own specific vocabulary for a theme while the labor reports use theirs, all while staying anchored to the exact same underlying idea. The result is a system where the computer knows that the "money" topic in one collection and the "wages" topic in another are not just similar, but are actually the same topic viewed through different lenses.
The researchers tested this approach on three massive collections of text spanning from 1922 to 2019. One collection contained historical American English from magazines and newspapers, another held articles from the Harvard Business Review, and the third consisted of reports from the International Labour Review. These sources cover very different worlds: one is a broad mix of general history, one focuses on corporate management, and the third deals with workers' rights and social policy. The team compared their new method against older techniques where the collections were analyzed separately and then matched up later. The difference was stark. When using the old method of separate analysis, the computer could only correctly identify matching themes between the collections about 18 percent of the time. In contrast, the new shared-map method achieved a success rate of nearly 98 percent. This means that almost every time the computer looked at a theme in the business magazines, it could instantly and accurately find the corresponding theme in the labor reports, even though the words used were completely different.
The study also showed that this new method did not sacrifice the quality of the analysis for the sake of alignment. The computer was still able to learn the specific nuances of each collection, capturing the unique vocabulary of business executives and the distinct language of labor organizers. By keeping the main map fixed and only allowing small adjustments, the system preserved the ability to track how a single theme, such as economic transformation, evolved over the decades. The researchers were able to trace a single thread of conversation about the economy as it shifted from the industrial crises of the 1930s, through the post-war expansion of corporate management, and into the digital age of the late twentieth century. In each era, the shared map showed how the business magazines focused on profit and stock markets, while the labor reports focused on wages and worker protections, yet both were clearly discussing the same historical moment.
This work suggests that the key to comparing different groups of texts over long periods is to treat the alignment of themes as a fundamental part of the learning process, rather than an afterthought. The researchers found that when they tried to fine-tune the entire system for each collection separately, the connection between the themes broke down, and the computer lost its ability to see the big picture. Their findings indicate that by building a stable, shared foundation first, it is possible to see how different communities talk about the same big ideas in their own unique ways. This approach offers a more reliable way for historians and social scientists to compare diverse sources of information, ensuring that when they trace a theme from one century to the next, they are truly following the same story, even if the words on the page have changed.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.