VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations
This paper introduces VetScore, a risk-weighted evaluation framework for veterinary long-form QA that assesses the faithfulness of generated claims to source excerpts while prioritizing claims based on their potential harm, achieving high alignment with expert judgments.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but instead of a crime scene, you are looking at a long, detailed story written by a very smart, but sometimes overconfident, robot. This robot is an Artificial Intelligence (AI) that has read thousands of books and can answer almost any question you ask. However, in the world of medicine—whether for humans or their furry, feathered, and scaly friends—getting a fact wrong isn't just a bad grade; it can be dangerous. If the robot says a dog needs a specific amount of medicine, but it got the number wrong, the dog could get sick.
To stop the robot from making up facts, scientists have taught it to "show its work" by adding citations, like little footnotes that point to the original books it read. It's like the robot saying, "I read this in a book!" But here's the tricky part: just because the robot points to a book doesn't mean it actually read the right page, or that it understood the sentence correctly. Sometimes, the robot might point to a book about cats when it's talking about dogs, or it might invent a sentence that sounds like it belongs in the book but isn't actually there. This is where the real detective work begins: we need a way to check not just if the robot cited a source, but if the story it told actually matches the source, and how bad it would be if the story was wrong.
This is exactly what the paper "VETSCORE" tackles. The researchers, working with veterinary experts, built a new tool to grade these AI-generated answers. They realized that not all mistakes are created equal. If the robot gets a date wrong (like saying a study happened in 2021 instead of 2022), it's annoying, but a dog probably won't get hurt. But if the robot gets the dosage of a heart medication wrong, that's a disaster. So, they created a system that doesn't just count errors; it weighs them. They break the AI's long answer down into tiny, bite-sized facts. Then, for each fact, they ask two questions: "Did the robot actually find this in the book it cited?" and "How much harm could this cause if it's wrong?" Finally, they combine these answers into a single score that tells us how safe and reliable the AI's advice really is.
The team tested this system by having AI models generate answers to real veterinary questions, then having human experts (veterinarians and students) grade the same answers. They found that their new "VETSCORE" tool was very good at matching the human experts' judgments, even when using smaller, faster AI models to do the grading. The results suggest that by focusing on the risk of a mistake rather than just the number of mistakes, we can build a much safer way to check AI in high-stakes fields like veterinary medicine. It's like having a safety inspector who doesn't just count how many bricks are loose in a wall, but specifically checks if the ones holding up the roof are secure.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.