Hallucinations in AI Chatbots: A Comparative Analysis
This paper investigates hallucinations in prominent AI chatbots like GPT-4, Gemini, Claude, and Copilot through literature review and surveys, highlighting their tendency to generate fabricated information and citation errors in academic contexts while proposing mitigation strategies such as manual review and fact-checking.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the last few years, a new kind of computer program has become a common tool for writing, coding, and finding information. These programs, known as large language models, are trained on vast amounts of text from the internet. They are designed to predict what word should come next in a sentence, allowing them to generate human-like responses to questions. While they are incredibly useful, they have a peculiar flaw: they sometimes state things that are completely false with total confidence. Researchers call these errors "hallucinations." A hallucination might look like a real fact, a genuine quote from a book, or a working line of code, but it is actually a fabrication. This is not just a minor glitch; when these programs are used for schoolwork, medical advice, or professional research, a made-up fact can lead to real-world confusion or danger. The question facing scientists and users alike is not whether these errors happen, but how often they occur, which programs are most prone to them, and how people can spot them before they cause trouble.
A team of researchers from the Usman Institute of Technology set out to understand this problem by looking at the four most popular AI chatbots available today: GPT-4, Gemini, Claude, and Microsoft Copilot. Instead of just testing the machines in a lab, the researchers asked 80 actual users about their real-world experiences. They wanted to know how often these users encountered false information, how they noticed the errors, and what they did to fix them. The study focused on tasks that require high accuracy, such as writing academic papers, writing computer code, and conducting research. The goal was to move beyond technical benchmarks and see how these tools behave in the hands of everyday people who rely on them for important work.
The results of the survey revealed that hallucinations are a persistent issue, even in the most advanced systems. Most of the people who took the survey said they use these AI tools every day or several times a day. They rely on them for coding, research, and writing. However, the majority also reported that they frequently encounter made-up information or fake references. The researchers found that these errors do not happen randomly; they tend to appear more often as a conversation gets longer. After a user has exchanged five to twenty messages with the bot, the likelihood of the program slipping up and providing inconsistent or false details increases. This suggests that the longer you talk to the machine, the more likely it is to lose its grip on the truth.
When users realized something was wrong, they did not rely on the computer to tell them. Instead, they used their own knowledge. The most common way people spotted a hallucination was when the AI's answer contradicted something they already knew to be true. Another major red flag was the presence of fake citations. If the program listed a book, a study, or a website that could not be found or did not exist, users knew immediately that the information was fabricated. The study showed that detecting these errors is still largely a human job. It depends on the user's ability to verify facts and their willingness to question the machine, rather than on any automatic safety switch built into the software itself.
The researchers also compared the four different chatbots to see if one was safer or more reliable than the others. They found that each system has its own strengths and weaknesses. GPT-4 was praised for its ability to reason through complex problems, while Gemini was noted for its strong ability to pull information from search results, which helps keep its answers grounded in reality. Claude stood out for its focus on safety and honesty, particularly in sensitive areas like health, where it showed high accuracy. Microsoft Copilot was recognized for its tight integration with other software tools, making it useful for productivity. However, none of the four models was perfect. All of them still produced false citations and unreliable code at times. The study concluded that while these tools are powerful, they are not yet trustworthy enough to be used without human oversight, especially in high-stakes fields like medicine or academic research.
So, what can be done to fix this? The survey participants offered a clear answer. They did not believe that the solution lies solely in making the AI smarter. Instead, they said the most effective way to reduce errors is to require the AI to show its work. The users overwhelmingly preferred responses that included verified citations, meaning the program must point to a real, existing source for every fact it claims. Human review was the second most popular solution, followed by automatic fact-checking systems. The people using these tools want transparency; they want to see where the information comes from so they can check it themselves. The researchers suggest that future improvements should focus on making the AI explain its sources and admit when it is unsure, rather than just trying to generate more text.
Ultimately, this study highlights a crucial reality about the current state of artificial intelligence. These tools are becoming more sophisticated and are being used for more complex tasks, but the problem of making things up has not gone away. The researchers found that users are aware of the risk and are actively working around it by double-checking facts and demanding evidence. For these systems to become truly reliable, especially in critical areas like healthcare and education, they must evolve to provide clear, verifiable sources for their answers. Until then, the responsibility for accuracy remains firmly in human hands, with the AI serving as a helpful assistant rather than a final authority.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.