CitiLink-Minutes: A Multilayer Annotated Dataset of Municipal Meeting Minutes
This paper introduces CitiLink-Minutes, a multilayer annotated dataset of over one million tokens from 120 European Portuguese municipal meeting minutes, featuring structured metadata, discussion subjects, and voting outcomes to advance research in Natural Language Processing and Information Retrieval for local governance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a city council meeting as a massive, chaotic library where hundreds of books are written every year. These books are the meeting minutes—official records of what politicians talked about, what laws they passed, and who voted for what.
For a long time, these books were locked in a dark room. They were written in a confusing style, full of legal jargon, and no one had ever organized them. Computer programs (AI) wanted to read them to help us understand how our cities work, but they couldn't because the books were too messy and there were no "indexes" or "table of contents" to guide them.
This paper introduces CitiLink-Minutes, a project that acts like a super-organized librarian for these city council books. Here is the simple breakdown:
1. The Problem: A Wall of Text
Think of city council minutes as a giant, unsorted pile of sticky notes.
- The Issue: If you want to know, "Did the city council vote to fix the potholes in the north district?" you have to read thousands of pages of text to find that one sentence.
- The Gap: While researchers have built great tools to analyze speech (like transcripts of TV debates), they haven't had good tools for written city records. They lacked a clean, labeled dataset to train their AI.
2. The Solution: The "CitiLink" Super-Index
The researchers created a new dataset called CitiLink-Minutes. Imagine they took 120 of these messy "sticky note" books from six different Portuguese cities and gave them a complete makeover.
They didn't just scan the pages; they added four layers of "highlighting" (annotations) to every single sentence, like a color-coded study guide:
- Layer 1: The "Who" (Personal Info): They found every name and role (like "Mayor" or "Councilor") and then erased the real names (replacing them with "Person A" or "Councilor B"). This is like blurring faces in a photo so the AI can learn the structure of the meeting without violating privacy.
- Layer 2: The "When & Where" (Metadata): They tagged the date, time, location, and meeting type. It's like putting a perfect library card on every book.
- Layer 3: The "What" (Subjects): They identified the main topics. Was the meeting about "Budgets"? "Roads"? "Schools"? They labeled every topic so you can search by theme.
- Layer 4: The "Result" (Voting): This is the most exciting part. They tracked exactly who voted "Yes," who voted "No," who stayed home, and what the final decision was. It turns a paragraph of text into a clear scoreboard.
3. The Human Touch
You might think a computer did this, but no. The researchers hired a team of human linguists (like expert editors) to read every single page and apply these tags by hand.
- Why? Because computers get confused by messy human language. A human can tell the difference between a joke and a serious vote; a computer often can't.
- The Result: They created a "Gold Standard" dataset. It's like a practice exam with the answer key included, so other researchers can train their AI to do the same job.
4. Testing the AI (The Baselines)
To prove this dataset works, the researchers tried to teach two types of AI to read these minutes:
- The "Reader" (Encoder): A model that reads text and extracts facts.
- The "Writer" (Generative): A model that tries to summarize or answer questions.
The Result: The "Reader" model (specifically a model trained on Portuguese) did a fantastic job. It could find voting results and topics with high accuracy. The "Writer" model was okay but made more mistakes. This tells us that for this specific job, a careful, fact-finding AI is better than a chatty, summarizing one.
5. Why Does This Matter?
Think of CitiLink-Minutes as opening the doors to the city council's "black box."
- For Journalists: They can instantly find how a specific politician voted on a specific issue.
- For Citizens: They can see exactly what their local government is doing without reading 50 pages of legal text.
- For Researchers: They finally have a clean, safe, and organized playground to build better tools for understanding democracy.
The Catch (Limitations)
The paper admits the dataset isn't perfect yet:
- It only covers six cities in Portugal. It's like having a map of only six towns; it might not work perfectly for a city in a different country with different rules.
- Some voting details are still a bit vague (e.g., we know a "party" voted, but not always every single person).
- The names are currently just "Person A," but future versions might use smarter ways to hide identities while keeping the data useful.
In a Nutshell
This paper is about taking the boring, messy, and hard-to-read records of city meetings and turning them into a clean, searchable, and privacy-safe database. It's the first time this has been done so thoroughly for written minutes, giving AI a chance to finally help us understand how our local governments really work.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.