← Latest papers
💻 computer science

Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining

This paper introduces a task-agnostic method for measuring training data influence by quantifying how much individual examples reduce the distance to final model parameters, revealing that literature-related data drive early pretraining while STEM data become more influential in later stages across various model configurations.

Original authors: Yuto Nishida, Hirokazu Kiyomaru, Yusuke Oda, Takashi Kodama, Chaoran Liu, Daisuke Kawahara, Yusuke Miyao, Max Müller-Eberstein, Masaru Isonuma

Published 2026-08-14
📖 4 min read☕ Coffee break read

Original authors: Yuto Nishida, Hirokazu Kiyomaru, Yusuke Oda, Takashi Kodama, Chaoran Liu, Daisuke Kawahara, Yusuke Miyao, Max Müller-Eberstein, Masaru Isonuma

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a super-smart robot how to speak like a human. You don't just give it one book; you dump the entire internet into its brain. This process is called "pretraining," and it's the foundation of the AI models we use today. But here's the tricky part: once the robot is finished learning, nobody really knows which specific sentences or stories were the most important. Did it learn to be good at math because of a specific Wikipedia article? Did it learn to write poetry because of a specific novel?

Scientists have been trying to answer this by asking, "If I remove this one sentence, does the robot get worse at a specific test, like solving math problems or writing a poem?" This is like trying to figure out which ingredient made a cake taste good by tasting the cake after you've already eaten it. The problem is, AI models are supposed to be good at everything, not just one test. If you only check how well they do on math, you might miss the fact that they learned how to be funny from a different part of the data. So, the big question is: Can we see how the robot's brain changes over time without forcing it to take a specific test?

This paper introduces a clever new way to watch the robot learn, without needing a specific test. Instead of asking, "Did this sentence help the robot pass a math quiz?", the researchers asked, "Did this sentence help the robot get closer to its final, finished brain?" They treated the robot's final state as a destination on a map. Every time the robot read a new sentence, it took a tiny step. The researchers measured whether that step moved the robot closer to the finish line or if it accidentally took a step backward.

They applied this method to 18 different versions of a famous AI model family, watching them learn through 154 different checkpoints (like snapshots of the robot's brain at different times). What they found was a fascinating story about how the robot's "diet" changes as it grows up.

In the very beginning, when the robot was just starting to learn, the most helpful sentences came from literature and stories. These were the "top contributors" that pushed the robot toward its final form. It was as if the robot needed to understand human stories and emotions first. However, as training continued into the middle and later stages, the story changed. The sentences that became most important were no longer stories; they were from STEM fields (Science, Technology, Engineering, and Math).

The researchers saw a clear "crossover." Early on, books and literature were the stars of the show. Later on, technical data about science and computers took over the spotlight. This suggests that the robot doesn't just learn "hard" things at the end; it actually needs a specific sequence of learning. It seems to need the foundation of human language and stories first, and only then does it really benefit from the complex, technical data to reach its full potential.

The authors also checked if their method was reliable. They compared their "distance-to-finish" measurement with the old "test-score" method and found that the old method gave very different answers depending on which test you picked. If you tested for math, you thought math data was important the whole time. If you tested for stories, you thought stories were important the whole time. Their new method, however, showed a consistent pattern across all the different models they tested, suggesting that this shift from stories to science is a real, fundamental part of how these AI models learn, not just an accident of how they were tested.

So, the main takeaway is that training an AI isn't just about feeding it a random mix of data. It's more like a curriculum. The paper suggests that the "best" data changes depending on how far along the AI is in its training. Early on, it needs stories; later on, it needs science. This gives us a new way to think about how to build better AI: maybe we should feed them stories first, and save the complex math for when they are ready for it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →