← Latest papers
🤖 AI

In-Context Examples Suppress Scientific Knowledge Recall in LLMs

This paper demonstrates that adding in-context examples to large language models can suppress their ability to recall and apply latent scientific knowledge, causing them to rely instead on empirical pattern fitting even when the examples are generated by the correct underlying formulas.

Original authors: Chaemin Jang, Woojin Park, Hyeok Yun, Dongman Lee, Jihee Kim

Published 2026-05-01
📖 5 min read🧠 Deep dive

Original authors: Chaemin Jang, Woojin Park, Hyeok Yun, Dongman Lee, Jihee Kim

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: When Examples Make You Forget What You Know

Imagine you are a brilliant student who has memorized the laws of physics, chemistry, and economics. You know exactly how to solve a problem using the correct formulas. But then, your teacher gives you a test with a twist: before you solve the problem, they show you 10 examples of similar problems with their answers.

You might think, "Great! These examples will help me remember the rules better."

This paper discovered the opposite is true. When the AI (Large Language Model) sees those examples, it actually stops using its memorized scientific laws and starts guessing based on patterns in the examples instead. It's like a master chef who, after seeing a few photos of a dish, forgets the recipe and just tries to guess the ingredients by looking at the pictures.

The authors call this "Knowledge Displacement." The examples don't reinforce the knowledge; they push it out of the way.


The Two Ways the AI Thinks

The paper describes two different "modes" the AI uses to solve scientific problems:

  1. The "Textbook Mode" (Knowledge-Driven):

    • How it works: The AI sees a problem (e.g., "How fast does this object cool down?"), recognizes the keywords, and pulls the correct formula from its memory (Newton's Law of Cooling). It calculates the answer step-by-step using math.
    • Analogy: This is like a mechanic who hears a strange engine noise, knows exactly which part is broken based on years of training, and fixes it using the right tool.
  2. The "Pattern Mode" (Example-Driven):

    • How it works: The AI ignores the formulas. Instead, it looks at the 10 examples provided in the prompt and tries to find a shortcut or a trend. It guesses the answer by saying, "In the examples, when the number went up, the answer went down, so I'll do the same."
    • Analogy: This is like a mechanic who has never studied engines but looks at 10 photos of cars with broken parts. They guess the fix by saying, "In all these photos, the tire was flat, so I'll just guess the tire is flat," without understanding why the engine is making noise.

The Discovery: When you give the AI examples (even correct ones), it almost always switches from "Textbook Mode" to "Pattern Mode." It stops doing real science and starts doing "pattern matching."


The Surprise: Sometimes This Helps, Sometimes It Hurts

The researchers tested this on 60 different science problems across five fields: Economics, Chemistry, Physics, Biology, and Geoscience. They found that while the AI always switched to the "Pattern Mode," the result on the final score was different for each subject.

Here is how it played out:

  • The "Hurt" Case (Economics & Chemistry):

    • What happened: The AI knew the formulas perfectly. When shown examples, it stopped using them and started guessing patterns.
    • Result: The score dropped significantly.
    • Analogy: Imagine a math genius who knows the multiplication table perfectly. If you show them a few examples of 2×2=42 \times 2 = 4, they might stop thinking "2 times 2" and start guessing "maybe it's 5?" just because the examples looked a certain way. They get the answer wrong because they forgot the rule they already knew.
  • The "Help" Case (Geoscience):

    • What happened: The AI knew the formula but was bad at remembering the specific numbers (like the density of rocks). It kept guessing the wrong numbers. When shown examples, it stopped trying to recall the numbers and just looked at the patterns in the examples to guess the answer.
    • Result: The score went up.
    • Analogy: Imagine a student who knows the formula for a recipe but always forgets how much salt to add. If you show them 10 photos of the finished dish, they might stop trying to remember the salt amount and just copy the taste from the photos. They get a better dish, but they didn't actually learn the recipe; they just got lucky with the pattern.
  • The "No Change" Case (Physics & Biology):

    • What happened: The AI switched modes, but the "Pattern Mode" happened to be just as good as the "Textbook Mode" for these specific problems.
    • Result: The score stayed the same.
    • Analogy: It's like driving a car. You can drive using the GPS (Textbook Mode) or by following a friend's car (Pattern Mode). If your friend is driving perfectly, you arrive at the same time. But you still aren't using the GPS anymore.

Why Does This Happen?

The paper found that the AI's "Textbook Mode" is very fragile. It relies on keywords.

  • If the problem says "Consumer Surplus," the AI thinks "Economics Formula."
  • If you change the words to "Alpha Wave" (a fake name for the same concept), the AI forgets the formula entirely.

However, the "Pattern Mode" doesn't care about words. It just looks at the numbers. So, when you give the AI examples with numbers, it immediately switches to the pattern mode because it's easier, even if the examples are perfectly consistent with the rules it already knows.

The Main Warning for Users

The paper concludes with a caution for anyone using AI for science:

Don't assume that giving the AI examples will help it use its knowledge.
In fact, it might do the opposite. It might make the AI stop thinking like a scientist and start thinking like a guesser. Sometimes this guesser gets lucky and gets a better score, but often it gets a worse score because it forgot the rules it was supposed to know.

In short: The examples didn't teach the AI anything new; they just made it forget what it already knew.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →