From inference to prediction: how machine learning is reconfiguring science
By analyzing 4.9 million publications, this study reveals how machine learning is reconfiguring scientific knowledge production by shifting the core-periphery structure of research and driving a two-wave displacement of traditional inferential methods by increasingly opaque predictive models, particularly in health and social sciences.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of science as a massive, bustling library where researchers are trying to solve mysteries. For decades, the librarians (scientists) had a golden rule: to solve a mystery, you had to explain how you found the answer. If you couldn't show your work, the answer didn't count. This was the era of "Inference," where the goal was to understand the inner workings of the universe, like figuring out exactly why a specific chemical reacts or how a disease spreads.
But recently, a new, super-smart robot assistant named Machine Learning (ML) has moved into the library. This robot is incredible at guessing the answers. It looks at millions of clues and says, "I'm 99% sure the answer is X!" The problem? The robot often can't explain why it thinks that. It just knows. This is the era of "Prediction."
This paper, which looked at a staggering 4.9 million scientific publications from 1990 to 2025, acts like a detective trying to figure out how this robot is changing the library. Here is what they found, without the jargon.
The Map of the Robot's Brain
The researchers built a giant map of the robot's knowledge. They found that the robot's "brain" has a core and a periphery, kind of like a city.
- The Core (Physical Sciences): The heart of the robot's brain is in Computer Science and Engineering. This is where the robot was built and where it lives. About 28% of all the robot's work happens here.
- The Main Neighborhood (Health Sciences): The robot is most popular in Medicine, which makes up 15% of the work. It's the biggest place where the robot is actually used to solve real-world problems.
- The Outskirts: Other fields like social sciences and life sciences are on the edges, slowly adopting the robot, but they aren't the main hubs yet.
The Great Swap: From "Why" to "What"
Here is the big twist the paper discovered. For a long time, scientists in medicine and social sciences relied on old-school tools like linear and logistic regression. Think of these as clear, glass-walled calculators. You could see every number and step inside them. If a doctor used one, they could say, "I know the patient has a disease because these three specific factors added up."
But starting around 2015, a new wave of tools called Deep Learning (and later, Large Language Models) started taking over. These are like black boxes. They are incredibly good at guessing the right answer, often better than the glass calculators, but you can't see inside them.
The paper found that in fields that used to demand transparency (like health and social sciences), the old glass calculators are shrinking. Since 2015, the use of these transparent tools has been growing slower than the rest of science, while the black boxes are exploding in popularity. In fact, in some fields, the old tools are actually losing ground.
The Two Waves of the Black Box
The paper suggests this change happened in two distinct waves, like two different types of fog rolling in.
Wave 1 (2015–2021): The "Do-It-Yourself" Fog
In this first wave, scientists built their own black boxes (like Convolutional Neural Networks). They were still in control. They could train the robot on their own data, write down how they did it in their papers, and other scientists could, in theory, check their work. It was opaque (hard to see inside), but the scientist was still the captain.
Wave 2 (Post-2022): The "Mystery Box" Fog
Then, after 2022, things changed. The fastest-growing tools are now Large Language Models (like the ones that write essays or chat with you). These aren't built by the scientists anymore; they are built by big companies.
- The Catch: Scientists can't see the training data. They can't see the code. They can't see the weights (the robot's "brain" settings). They just type a question (a "prompt") and get an answer.
- The Result: The paper suggests this creates a new kind of "intellectual debt." Scientists are using systems whose logic is completely unknown to them. The control has shifted from the scientist to the company that owns the model.
What This Means for Science
The paper argues that this isn't just a technical upgrade; it's a change in how science decides what is "true."
- Before: Truth was about understanding. "Show me the steps, and I'll believe you."
- Now: Truth is increasingly about accuracy. "Give me the right answer, even if I can't see how you got it."
The authors point out that this creates a tension. Science has always promised to be open and transparent (the "Open Science" ideal), but these new tools are often secret and proprietary. The paper suggests that while science is getting better at predicting things, it might be losing its ability to explain them, especially in fields where understanding the "why" has always been the most important part.
The Bottom Line
The paper doesn't say this is a disaster, but it does say it's a massive shift. We are moving from a world where scientists explain the rules of the game to a world where they just use a powerful, mysterious tool to win the game. The paper suggests that while this expands what science can do, it fundamentally changes how knowledge is made, validated, and trusted. The old glass calculators aren't gone, but in many places, they are being replaced by black boxes that work better, but keep their secrets.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.