Scaling In, Not Up? Testing Thick Citation Context Analysis with GPT-5 and Fragile Prompts
This paper demonstrates that while GPT-5 can serve as a stable co-analyst for thick citation context analysis by generating diverse interpretative hypotheses, its specific readings and vocabulary are systematically shaped by prompt scaffolding, often shifting focus from admonishment to lineage compared to human expert reconstructions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand a secret handshake between two people. You see Person A mention Person B in a speech. On the surface, it looks like a simple "Thank you." But is it? Or is it a subtle dig? A way of saying, "I'm smarter than you," or "I'm borrowing your idea without giving you full credit"?
For decades, scientists have tried to use computers to sort these "handshakes" (citations) into neat categories like "Praise," "Criticism," or "Just a Reference." They wanted to scale up: build a robot that can read a million papers and label every single one quickly.
This paper asks a different question: What if we stop trying to force the robot to give us one single label? What if, instead, we ask the robot to scale in? What if we ask it to sit down, read the same tricky sentence over and over again, and generate many different, plausible stories about what that sentence might mean?
Here is the story of the experiment, explained simply.
The "Hard Case": A Tiny Footnote
The author picked a famous, tricky little footnote from a 1975 science paper. It's like a tiny, cryptic note in the margin of a book.
- The Text: "We borrowed this idea from Person X, then Person Y repeated it. We'll stop worrying about the difference between them for now."
- The Mystery: A famous scholar named Gilbert looked at this years ago and said, "Wait a minute. Person Y repeated Person X, but they didn't actually credit Person X when they defined the idea! This footnote is actually a polite, hidden insult. It's saying, 'Hey, you forgot to give credit, so we are going to fix the history books for you.'"
The author wanted to see if a super-smart AI (GPT-5) could figure this out, not by guessing one answer, but by acting like a detective who writes down every possible theory.
The Experiment: The "Two-Stage Detective"
The author didn't just ask the AI, "What does this mean?" That's too vague. Instead, they built a two-stage pipeline:
- Stage 1: The Surface Scan.
The AI looked only at the footnote. It acted like a standard robot, saying, "This looks like a normal 'Thank You' note." (It was very consistent here, just like a human would be). - Stage 2: The Deep Dive.
Now, the AI got the whole story. It saw the original paper, the paper it cited, and the paper that cited it. It was told: "Okay, you said it's a 'Thank You' note. But now, look at the clues. Are there hidden meanings? Write down five different theories about what is really happening."
The Twist: The "Prompt Nudge"
Here is where it gets really interesting. The author realized that the AI's "personality" changes based on how you talk to it. So, they tried three different ways of asking the question:
- The "Neutral" Ask: "Just tell me what you think."
- The "Toward" Nudge: The author gave the AI a little hint, like a teacher whispering, "Remember, sometimes people use citations to rewrite history or hide a critique."
- The "Away" Nudge: The author gave a different hint: "Sometimes people use citations just to make their list look long, or to teach students."
What Happened?
The results were fascinating, like watching a chameleon change colors based on the background.
The AI is a Great "Idea Generator":
The AI didn't just give one answer. It produced 450 different theories. Some were boring ("They just wanted to be polite"). Some were brilliant ("They were trying to silence a rival"). Some were wild ("They were trying to teach a student").- Analogy: Imagine asking a friend to guess what a stranger is thinking. A normal person might say, "He looks happy." The AI, when prompted right, says, "He looks happy, OR he's hiding a secret, OR he's nervous about a test, OR he's pretending to be happy to impress someone." It widens the circle of possibilities.
The "Nudge" Changed the Story:
- When the author nudged "Toward" (hinting at hidden critiques), the AI started writing more stories about power struggles, credit-stealing, and hidden insults. It started sounding like the famous scholar Gilbert.
- When the author nudged "Away" (hinting at teaching or social politeness), the AI stopped looking for fights and started writing stories about teaching students or building bridges between groups.
- Analogy: It's like asking a detective to solve a crime. If you say, "Think like a spy," the detective finds spies. If you say, "Think like a teacher," the detective finds a classroom. The clues are the same, but the story changes based on the hint you give.
The AI Missed the "Big Insult":
Even with the hints, the AI rarely concluded that the footnote was a direct insult. It preferred to think the authors were just "managing their reputation" or "organizing the family tree of ideas." It was a bit too polite to imagine a full-blown fight.- Why? The AI is trained to be helpful and nice. It's like a diplomat who sees a tense meeting and thinks, "Oh, they are just negotiating," rather than "They are about to start a war."
The Big Takeaway
This paper isn't about replacing human scientists with robots. It's about changing how we use robots.
- Old Way (Scaling Up): "Robot, read 1 million papers and tell me if this is 'Good' or 'Bad'." (The robot often gets it wrong because human meaning is messy).
- New Way (Scaling In): "Robot, read this one tricky sentence. Act as my co-analyst. Give me 10 different ways to interpret it. Show me the evidence for each. Then, I (the human) will decide which story makes the most sense."
The Final Metaphor:
Think of the AI not as a judge (who gives a final verdict), but as a creative writing partner.
If you ask a writing partner, "Write a story about this sentence," they might give you a tragedy, a comedy, and a mystery. They don't tell you which one is "true." They just hand you a menu of possibilities. Then, you, the human expert, pick the one that fits the real world.
The paper shows that if we treat AI as a tool to expand our imagination rather than just a tool to speed up our labeling, we can uncover hidden meanings in science that we might have missed on our own. But we have to be careful: the way we ask the question changes the answer, so we must always check the robot's work.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.