Generative AI and the future of scientometrics: current topics and future questions
This paper proposes a conceptual framework distinguishing between the semantic and pragmatic dimensions of text to guide the principled, explainable integration of generative AI in scientometrics, while warning of its potential to alter the fundamental textual characteristics used to measure science.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine scientometrics as the "weather forecasting" of the academic world. Instead of measuring rain and wind, these scientists measure words, citations, and author names to understand how knowledge grows and moves.
Now, imagine Generative AI (GenAI) as a super-fast, incredibly well-read robot librarian who has read almost every book ever written. This paper asks two big questions:
- Can we trust this robot librarian to help us do our weather forecasting?
- What happens to our weather maps if the robot librarian starts writing half the books in the library itself?
Here is the breakdown of the paper's arguments using simple analogies.
1. The Robot's Skill Set: The "Three-Layer Cake"
The authors explain that AI isn't magic; it works by predicting the next word in a sentence based on patterns it learned. To understand where the robot is good and where it fails, they use a "three-layer cake" model of language:
- Layer 1: Syntax (The Grammar)
- The Analogy: Building a perfect Lego tower.
- The Robot's Skill: Excellent. The robot is a master builder. It knows exactly how to stack words so they look grammatically correct and follow the rules. If you ask it to write a sentence, it will almost never make a grammar mistake.
- Layer 2: Semantics (The Meaning)
- The Analogy: Understanding what the Lego tower represents (e.g., "This is a castle").
- The Robot's Skill: Very Good. The robot is great at understanding general meanings. If you ask it to sort a pile of books into "History" or "Science," it does a great job because it recognizes the patterns of words used in those fields. It can handle tricky words that have different meanings depending on the context (like the word "bank" meaning a river edge or a place for money).
- Layer 3: Pragmatics (The Context & Purpose)
- The Analogy: Knowing why someone built the tower and who they are building it for. Is it a gift? A joke? A serious architectural plan?
- The Robot's Skill: Weak and Unpredictable. This is where the robot stumbles. It doesn't truly understand the social situation, the hidden jokes, or the specific human intent behind a text. It can guess, but it often misses the mark when judging individual people or specific, complex situations.
2. Using the Robot for Science (The "Tool" Section)
The authors tested how well this robot librarian helps with scientometrics tasks.
- Where the Robot Shines (Semantic Tasks):
If you need to sort thousands of papers into broad categories (like "Biology" vs. "Physics") or summarize a huge list of topics, the robot is fantastic. It's like having a machine that can instantly sort a mountain of mail into the right bins. It's fast and consistent for big groups. - Where the Robot Struggles (Pragmatic Tasks):
If you ask the robot to judge the quality of a single paper, decide if a specific scientist is "prominent," or predict exactly how many citations one specific article will get, the robot is unreliable. It's like asking a robot to judge a stand-up comedy routine; it might get the general vibe, but it will miss the specific nuances that make a human laugh or cry.- The Paper's Advice: Don't let the robot make the final decision on individual cases. Use it to do the heavy lifting (sorting, summarizing), but keep a human expert in the loop to check the work, especially for high-stakes decisions like funding or hiring.
3. The Danger: The Robot Becoming the Author (The "Object" Section)
This is the paper's most worrying and exciting part. So far, we've talked about the robot helping us analyze science. But what if the robot starts writing the science?
The authors warn that if scientists start using AI to write their papers, the "weather maps" of science might get distorted. Here is how:
- The Vocabulary Shift: AI tends to write in a very specific, polished style. If everyone uses it, scientific writing might become "homogenized" (all sounding the same). This could make it look like research topics are more stable than they really are, hiding new, weird, or emerging ideas.
- The Authorship Blur: If an AI helps write a paper, who is the author? The paper argues that this creates a gap between the name on the paper and the actual human effort. It might look like one person is doing a massive amount of work, when in reality, they are just using a tool to speed things up.
- The Reference Trap: AI sometimes "hallucinates" (makes up) fake references. Even if it stops doing that, it might start citing the same popular papers over and over because that's what its training data suggests. This could make the "map" of science look like a few stars are shining much brighter than they actually are, while new, lesser-known ideas get ignored.
The Bottom Line
The paper concludes that Generative AI is not just a new tool; it is changing the very fabric of the data scientists use to study science.
- For the Robot: We should use it for big-picture sorting and summarizing (Semantics), but we must be very careful about using it for judging individual people or specific contexts (Pragmatics).
- For the Future: We need to be vigilant. If AI starts writing our papers, the "metrics" we use to measure scientific success (like citation counts or word usage) might stop telling the truth about how science is actually progressing.
The authors call for a new era of research: we need to study how AI is changing the language of science, so we don't lose our ability to understand the real story behind the data.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.