← Latest papers
💬 NLP

DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification

This paper presents the DS@GT ARC system for CheckThat! 2026, which demonstrates that an LLM-based verifier with Best-of-N selection outperforms a lightweight TF-IDF reward model in multilingual numerical claim verification, while also showing that AraBERT surpasses multilingual baselines for Arabic and that sub-claim decomposition fails to improve performance.

Original authors: Sagnik Sinha, Shreyas Shrestha

Published 2026-07-29
📖 4 min read☕ Coffee break read

Original authors: Sagnik Sinha, Shreyas Shrestha

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, chaotic library where anyone can write a book, and sometimes those books are filled with lies. In this digital age, false stories spread faster than a rumor in a high school hallway, often hiding behind a wall of numbers and dates that make them look super serious. This is the world of "fact-checking," a field where scientists try to build robot detectives to sort truth from fiction. But here's the tricky part: when a lie includes a specific number—like "50% of people did this"—it feels more real to our brains, even if it's made up. This is called the "numeric truth effect." To catch these tricky lies, we need computers that don't just read words, but also do math and check dates, a skill that has been notoriously hard for machines to master.

Now, picture a team of researchers from Georgia Tech entering a high-stakes competition called "CheckThat!" to build the ultimate fact-checking robot. Their mission? To create a system that can verify claims about numbers and dates in both English and Arabic. But they didn't just want a robot that guesses "True" or "False." They wanted a robot that could look at a pile of different reasoning paths (like different detectives writing up their case files) and pick the best one to decide the verdict. Think of it like a judge who doesn't just want a verdict, but wants to see the best-written legal argument to be sure the decision is solid.

The team, known as DS@GT, tried two very different strategies to solve this puzzle. The first approach was like hiring a super-smart, well-read detective (a Large Language Model) and giving it a special training session to learn how to grade other detectives' work. They used a technique called "LoRA," which is like giving the detective a set of specialized highlighters to focus on the most important parts of the case without rewriting their whole brain. They even tried breaking big, complicated claims into smaller, bite-sized pieces first, hoping it would make the job easier.

The second approach was more like a quick, clever librarian. Instead of a super-intelligent detective, they built a lightweight tool that looked for specific patterns, like counting how many numbers or dates matched between the claim, the evidence, and the reasoning. It was a "reward model" that gave points for matching numbers and penalized mismatches, then grouped the results to see which verdict had the most support.

So, what did they find? The "super-smart detective" (the LLM approach) turned out to be the overall winner, especially at finding the right reasoning path quickly. It was better at spotting the correct answer in most situations. However, the "quick librarian" (the reward model) had a secret superpower: it was surprisingly good at handling the messy, confusing cases where the evidence was split between "True" and "False," a category the researchers call "Conflicting."

Interestingly, the team's idea to break big claims into smaller pieces didn't work as hoped. Instead of making the job easier, it actually added noise, like trying to solve a puzzle by cutting the pieces into even smaller, confusing fragments. For the Arabic language part of the challenge, they found that using a model specifically trained on Arabic (AraBERT) was much better than using a general model that speaks many languages, proving that sometimes a specialist is better than a generalist.

In the end, the paper suggests that while big, smart AI models are currently the best all-around fact-checkers, we shouldn't ignore the value of simpler, pattern-matching tools when dealing with the most ambiguous and tricky claims. They didn't solve the problem of misinformation forever, but they gave us a clearer map of which tools work best for which kind of lie.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →