← Latest papers
💬 NLP

CitiLink-Summ: Summarization of Discussion Subjects in European Portuguese Municipal Meeting Minutes

This paper introduces CitiLink-Summ, the first European Portuguese corpus of municipal meeting minutes with 2,322 manually crafted summaries, to address the scarcity of data for automatic summarization in this domain and establish baseline performance for state-of-the-art generative models and large language models.

Original authors: Miguel Marques, Ana Luísa Fernandes, Ana Filipa Pacheco, Rute Rebouças, Inês Cantante, José Isidro, Luís Filipe Cunha, Alípio Jorge, Nuno Guimarães, Sérgio Nunes, António Leal, Purificação Silvano, Ri
Published 2026-02-19
📖 4 min read☕ Coffee break read

Original authors: Miguel Marques, Ana Luísa Fernandes, Ana Filipa Pacheco, Rute Rebouças, Inês Cantante, José Isidro, Luís Filipe Cunha, Alípio Jorge, Nuno Guimarães, Sérgio Nunes, António Leal, Purificação Silvano, Ricardo Campos

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you walk into a city council meeting. The room is packed, the air is thick with formal language, and the officials are discussing dozens of different topics: fixing a pothole, approving a new park, debating a budget. By the time the meeting ends, the official record (the "minutes") is a massive, 50-page document written in stiff, legalistic Portuguese.

For the average citizen, trying to read that document to find out what happened is like trying to find a specific needle in a haystack that is on fire. It's too long, too boring, and too confusing.

Enter "CitiLink-Summ."

This paper introduces a new tool and a new "training gym" for computers to solve this problem. Here is the simple breakdown of what the researchers did:

1. The Problem: The "Wall of Text"

Municipal meeting minutes are essential for transparency, but they are terrible for readability. They are dense, formal, and overwhelming. If you want to know if your city is building a new library, you shouldn't have to read 40 pages of legal jargon to find that one sentence.

2. The Solution: A New "Gym" for AI

To teach computers how to summarize these documents, you need a teacher (a dataset). But for European Portuguese, there was no textbook. It was like trying to teach someone to swim in a desert because there was no pool.

The researchers built that pool. They created CitiLink-Summ, a massive dataset containing:

  • 120 real meeting documents from six different Portuguese cities.
  • 2,880 human-written summaries.

How did they make it?
Imagine a team of four expert linguists (the "teachers") sitting down with these documents. They didn't just copy-paste sentences. They read a specific topic, understood the core idea, and then rewrote it in their own words as a short, clear summary.

  • Analogy: If the original text was a 20-minute rambling speech, the summary was a crisp 30-second news clip.
  • They did this with extreme care, checking and re-checking to ensure the summaries were accurate but also abstract (meaning they didn't just copy words; they truly understood the meaning).

3. The Challenge: "Abstract" vs. "Copy-Paste"

The researchers wanted to make sure the summaries were hard for computers to fake. They measured how much the summaries "reused" words from the original text.

  • Low Reuse (High Abstraction): The computer has to think and rephrase.
  • High Reuse (Low Abstraction): The computer just cuts and pastes.

They found that their human summaries were highly abstract. This is good news because it means the dataset is a tough, high-quality test. It forces the AI to actually learn to summarize, not just to be a photocopier.

4. The Test Drive: Putting AI to Work

Once they built the dataset, they put it to the test. They took the smartest AI models available (like BART, PRIMERA, and even the giant Gemini) and asked them to summarize the documents.

  • The Results: The AI models did a decent job, but they weren't perfect yet. The best models (the "big brains" like PRIMERA) got the highest scores, but there is still a lot of room for improvement.
  • The Takeaway: It's like showing a student a math problem. They got the right answer most of the time, but they still make silly mistakes. This dataset gives researchers a clear "answer key" to help them train the AI to get better.

5. Why Does This Matter?

This isn't just about fancy computer science; it's about democracy.

  • For Citizens: Imagine an app where you can type, "What did the Porto council decide about the new bus route?" and get a clear, 2-sentence answer instantly.
  • For Transparency: It makes government less of a secret club and more of an open conversation.

The Bottom Line

The authors have built the first-ever "training manual" for summarizing Portuguese city council meetings. They've shown that while AI is getting good at this, it still needs more practice. By releasing this data to the public, they are inviting other scientists to help build the ultimate "City Council Translator" that will make local government accessible to everyone, not just lawyers and experts.

In short: They turned a boring, unreadable wall of text into a clear, teachable lesson for computers, with the ultimate goal of helping regular people understand what their city is doing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →