Finetuning with Scientific Data Increases Hallucinations: A Multi-domain Factuality Evaluation of LLMs
This paper introduces SciFactCheck, a multi-domain benchmark revealing that scientific fine-tuning paradoxically increases hallucinations and assertiveness in large language models while decreasing their internal confidence, thereby challenging current domain-specific fine-tuning approaches for improving factuality.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a brilliant, well-read librarian who knows a little bit about everything. This librarian is your standard Large Language Model (LLM). They are great at chatting, but sometimes they make things up when they aren't sure of the facts. This is called "hallucinating."
Now, imagine you take that librarian and send them to a specialized university to study only science. You give them thousands of textbooks, research papers, and journals. You expect them to come back as a "Scientific Expert" who is even more reliable than before.
This paper is a report card on that experiment, and the news is surprisingly bad.
Here is what the researchers found, using simple analogies:
1. The "Specialist" Got Worse at Truth
The researchers created a massive test called SCIFACTCHECK. They asked 18 different AI models (some general, some "scientifically trained") to write short paragraphs about 2,500 different scientific topics, from engineering to philosophy.
The Result: The "Scientific" models actually made more mistakes than the general ones.
- The Analogy: It's like taking a general chef and sending them to culinary school to learn only about spices. You'd expect them to be better at cooking, but instead, they started forgetting how to boil water and over-seasoned every dish. The specialized training seemed to confuse the models, making them less reliable on basic facts than their general-purpose cousins.
2. The Three Types of "Lying"
The study didn't just look at "wrong" answers; it looked at how they were wrong. They found three specific ways the models hallucinated:
- The "Unverifiable" Lie (Unverifiability): The model makes up a fact that sounds real but can't be found in any book or paper.
- Analogy: The librarian says, "In 1995, a secret society of cats ruled the internet." It sounds specific, but no record of this exists.
- The "Overconfident" Lie (Overclaim): The model takes a small truth and blows it up into a huge, absolute certainty.
- Analogy: A study says, "This drug might help some patients." The model says, "This drug cures everyone." It exaggerates the scope and certainty.
- The "Fake Citation" Lie (Attribution): The model invents a source to back up its claim.
- Analogy: The librarian says, "According to the famous 'Journal of Space Physics' article by Dr. Smith..." but that article and that doctor never existed.
3. The "Confident but Clueless" Paradox
One of the strangest findings was about how the models felt about their own answers.
- The Finding: The scientifically trained models used more confident language (words like "definitely," "proven," "always") but were actually less confident internally (their computer "gut feeling" was shaky).
- The Analogy: Imagine a student who doesn't know the answer to a math problem. Instead of saying, "I'm not sure," they stand up, raise their hand, and shout the answer with 100% conviction. They have learned the style of a confident scientist, but they haven't learned the substance. They are "loudly wrong."
4. The "Fact-Checker" Problem
Finally, the researchers tried to use computers to grade these models, and then asked human experts to grade them too.
- The Result: The computer graders and the human experts didn't always agree. Even the humans struggled to agree on what counted as a "scientific fact" that could be checked.
- The Analogy: It's like having a group of judges try to score a gymnastics routine, but they can't even agree on the rules of the sport. In some fields (like math), everyone agrees on the rules. In others (like philosophy or social science), it's very hard to define what is "checkable."
The Bottom Line
The paper concludes that simply training AI on scientific books does not make it a trustworthy scientific expert. In fact, it might make it more dangerous because it sounds more confident while being less accurate. We need better tools to verify scientific claims, and we can't just assume that "specialized" AI is automatically better than "general" AI.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.