← Latest papers
💬 NLP

ScienceMeter: Tracking Scientific Knowledge Updates in Language Models

This paper introduces ScienceMeter, a framework for evaluating scientific knowledge updates in Large Language Models through three metrics—preservation, acquisition, and projection—revealing that current update methods struggle to maintain existing knowledge while effectively learning new facts and anticipating future scientific advancements.

Original authors: Yike Wang, Shangbin Feng, Yulia Tsvetkov, Hannaneh Hajishirzi

Published 2026-07-22✓ Author reviewed
📖 4 min read☕ Coffee break read

Original authors: Yike Wang, Shangbin Feng, Yulia Tsvetkov, Hannaneh Hajishirzi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot librarian who has read almost every book ever written up to a certain year. This librarian is amazing at answering questions about history, science, and pop culture from that era. But here's the catch: the world doesn't stop turning. New scientific papers are published every single day, revealing fresh discoveries that the librarian hasn't seen yet. If you ask this robot about a breakthrough that happened yesterday, it might confidently tell you something wrong, or it might just say, "I don't know," because its brain is frozen in the past. This is the problem of "stale knowledge" for Artificial Intelligence. Scientists are trying to figure out how to feed these robots new information without making them forget what they already knew or causing them to hallucinate wild, incorrect facts. It's like trying to teach a student a new chapter in a textbook without them forgetting the previous chapters or making up their own fake facts to fill in the gaps.

Enter ScienceMeter, a new framework created by researchers at the University of Washington to test exactly how well we can update these AI brains with fresh scientific knowledge. Think of the researchers as "knowledge mechanics" trying to tune up a high-performance engine. They didn't just ask, "Did the robot learn the new stuff?" Instead, they built a three-part stress test to see if the robot could:

  1. Remember the old stuff (Preservation): Did it forget the basics it learned years ago?
  2. Learn the new stuff (Acquisition): Did it actually understand the new papers it was just shown?
  3. Guess the future stuff (Projection): Could it use what it just learned to make a smart guess about a discovery that hasn't happened yet?

To run this test, they created a massive dataset covering ten different scientific fields, from computer science to medicine. They treated scientific knowledge like a timeline: "Past" papers (what the robot already knew), "New" papers (what they tried to teach it), and "Future" papers (what they hoped the robot could predict). They then tried five different methods to update the robot's brain, ranging from simply reading the new papers during a chat (inference) to actually retraining the robot's brain with the new data (training).

The results were a bit of a reality check. Even the best methods struggled to do all three things at once. The top-performing update method managed to keep 85.9% of the old knowledge safe, successfully learned 71.7% of the new knowledge, but could only project 37.7% of the future knowledge. In other words, the robot was pretty good at remembering the past, decent at learning the present, but still very shaky at predicting the future.

The study also found that the size of the robot mattered. Bigger, more powerful models were surprisingly good at learning new things just by reading them during a conversation (inference), almost like a genius student who can learn a new subject just by reading a textbook once. However, smaller models needed to be "retrained" from scratch to learn the new stuff effectively; just showing them the new info wasn't enough.

Perhaps the most interesting finding was that the type of science mattered. In fields where knowledge changes slowly, like political science, the robots did great at remembering and predicting. But in fast-moving fields like materials science, where new discoveries happen constantly, the robots got confused much faster. The researchers suggest that the more volatile a field is, the harder it is for an AI to keep its knowledge up to date without mixing things up.

Ultimately, the paper concludes that while we are getting better at updating AI, we haven't cracked the code yet. There is no magic button that lets an AI instantly learn new science, keep all its old facts, and predict the future all at the same time. The researchers argue that developing a method that can truly balance these three goals is one of the biggest challenges left in making AI a reliable partner for real scientific discovery. They warn that if we can't solve this, our AI tools might become confident experts in yesterday's science while being completely clueless about tomorrow's breakthroughs.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →