On the Role of Citations in Preference Data
This paper investigates how human judges and open-source LLMs evaluate citations in scientific question answering, revealing that humans prefer diverse but fewer citations while LLMs exhibit citation-related preferences that vary by model and data, thereby offering critical insights for improving preference data collection in reward modeling.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving world of artificial intelligence, a specific type of task has become central to how these systems learn and improve: asking a computer to choose the better of two answers. This process, known as preference learning, is the engine behind modern AI training. When a system generates two different responses to a question, a judge—either a human expert or another AI program—decides which one is superior. These choices are then used to teach the original system how to behave more like a trusted expert. However, a critical piece of the puzzle has remained unclear: when these judges look at answers that include references to outside sources, known as citations, what exactly are they looking for? Do they value the sheer number of sources, the variety of those sources, or the freshness of the information? Understanding this is vital because citations are meant to be a shield against fabrication, allowing users to verify that an AI is telling the truth. If the systems that train AI do not understand how to weigh these references correctly, the resulting models may learn to prioritize the wrong things, such as listing many sources without checking their quality, or ignoring the credibility of the information entirely.
A team of researchers set out to investigate this exact question by examining how both human experts and artificial intelligence programs evaluate scientific answers that include citations. They focused on a scenario where an AI is asked a complex scientific question and provides a long, detailed response backed by references to academic papers. The researchers gathered thousands of these pairs of responses, where one answer was chosen over the other by human experts, and then asked four different large language models to make the same choices. To understand the hidden rules behind these decisions, the team used a statistical method that allowed them to isolate specific features of the answers, such as how many citations were used, how diverse the sources were, and whether the references matched the text, while controlling for other factors like the length of the answer.
The study revealed a distinct and somewhat surprising pattern in how humans judge these responses. When people reviewed the scientific answers, they preferred responses that drew from a wide variety of different sources, suggesting they valued breadth and perspective. However, they also preferred answers that had fewer citations overall. This indicates that while humans want to see that an answer is well-researched, they are put off by responses that feel cluttered with an excessive number of references. Furthermore, humans were quick to penalize answers where the in-text references did not match the list of sources at the end, a common error that suggests the AI might be making up references. Interestingly, the age of the sources mattered very little to human judges; they did not show a strong preference for newer papers over older ones.
In contrast, the artificial intelligence judges behaved quite differently, and their preferences varied depending on which specific model was doing the judging. While the AI models also tended to like diverse sources, they often showed a much stronger reaction to the number of citations than humans did. Some of the AI judges strongly disliked answers with many citations, while others were indifferent. More critically, the AI models failed to notice or penalize the mismatched references that humans found so problematic. Instead, the AI judges were heavily influenced by superficial features, such as the length of the response. Several of the models consistently preferred longer answers, regardless of whether the extra length added value or just filled space. This suggests that without specific training or access to the actual source documents, these AI systems are not truly "reading" the citations to verify their accuracy but are instead reacting to the visual or structural patterns of the text.
The researchers found that these differences were not just minor quirks but significant divergences that could affect how future AI systems are trained. Because AI models are often used to label data for training other models, relying on them without understanding their specific biases could lead to a system that learns to prioritize length or citation count over actual accuracy and source diversity. The study highlights that while AI can mimic human judgment in many ways, it lacks the intuitive understanding of what makes a reference credible or consistent. The authors suggest that to build better systems, developers must be more careful about how they collect preference data, ensuring that human guidelines are clear about citation quality and that AI judges are either given tools to verify sources or are selected based on a deep understanding of their specific biases. Ultimately, the work serves as a reminder that in the quest to make AI more reliable, the details of how we ask it to choose between answers matter just as much as the answers themselves.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.