ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles
This paper introduces ClaimPT, a high-quality European Portuguese dataset of 1,308 news articles from LUSA News Agency containing 6,875 annotated factual claims, designed to address the scarcity of resources for automated fact-checking in low-resource languages and establish baseline benchmarks for claim detection.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a massive, chaotic marketplace. Every day, millions of people shout out facts, opinions, rumors, and lies. Fact-checking is the job of the "truth inspectors" who run around trying to verify if what people are shouting is actually true.
The problem? There are too many shouts, and the inspectors are too few. By the time an inspector verifies a lie, it has already traveled around the world three times. We need robots to help, but robots are currently very bad at this job because they haven't been trained on enough examples, especially in languages other than English.
This paper introduces ClaimPT, a new tool designed to teach these robots how to spot the truth in Portuguese news. Here is a simple breakdown of what they did:
1. The Problem: The "Needle in a Haystack"
In the world of news, most sentences are just background noise (like "The meeting started at 9 AM"). But hidden inside are specific claims—statements like "The government promised to build a bridge." These are the "needles" that need to be checked.
Finding these needles is hard because:
- They are often buried in long articles.
- They are mixed with opinions and stories.
- Until now, there was no "training manual" for computers to learn how to find them in Portuguese. Most training data was in English, like teaching a dog to fetch only using English commands.
2. The Solution: Building a "Training Gym" (ClaimPT)
The authors built a massive training gym for AI. They partnered with LUSA, a major Portuguese news agency, to get 1,308 real news articles.
Think of this dataset as a highlight reel for a sports coach.
- The Players: The news articles.
- The Coaches: Two human experts who read every single sentence.
- The Referee: A third expert who checked the coaches' work to make sure they were fair.
These humans went through the articles and put "stickers" on the text. They marked:
- The Claim: The specific sentence that makes a verifiable fact.
- The Source: Who said it? (e.g., "The Mayor" or "A local resident").
- The Time: When did they say it?
- The Topic: Is this about politics, health, or sports?
They didn't just mark the sentence; they also marked who said it and what they were talking about, creating a rich, detailed map of the truth.
3. The Challenge: Teaching the Robot
Once the gym was built, the authors tried to teach a robot (an AI model) to do the job. They tested two types of robots:
- The "Generative" Robot: A fancy AI that tries to write out the answer from scratch (like a student taking a test without notes).
- The "Encoder" Robot: A specialized AI trained to spot patterns (like a detective looking for fingerprints).
The Result:
The "Generative" robot got confused. It often thought that any quote in a newspaper was a claim, even if it was just an opinion. It was like a security guard who stops everyone who speaks, rather than just the people breaking the law.
The "Encoder" robot did much better. It learned to distinguish between a hard fact ("The bridge was built") and a soft opinion ("I think the bridge is ugly"). However, even the best robot still struggled a bit. This is because finding a specific claim inside a long news story is like finding a specific word in a book without knowing which page it's on.
4. Why This Matters
Before this paper, if you wanted to build a fact-checking app for Portuguese news, you were flying blind. You had no map.
ClaimPT is that map.
- For Researchers: It gives them a standard test to see how good their new AI models are.
- For the Public: It helps build tools that can automatically flag fake news or verify political promises in Portuguese, making the information ecosystem healthier.
The Big Picture
Think of this paper as the blueprint and the first batch of bricks for a new factory. The factory's job is to automatically sort truth from fiction in Portuguese news. The authors haven't built the whole factory yet, but they have provided the essential materials (the dataset) and the instructions (the guidelines) so that engineers around the world can start building it.
In short: They took a messy, unorganized pile of Portuguese news, carefully labeled the "truthful statements," and handed the labeled pile to the world so computers can finally learn how to fact-check in Portuguese.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.