Can Commercial LLMs Be Parliamentary Political Companions? Comparing LLM Reasoning Against Romanian Legislative Expuneri de Motive
This paper evaluates six commercial large language models against Romanian legislative reasoning, finding that while frontier models achieve high semantic accuracy on standardized tasks, all models exhibit task-dependent confabulation and contextual ignorance, suggesting that the primary risk for legislators is not ideological bias but the compounding effects of bounded rationality across the human-AI advisory chain.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Question: Can AI Be a Politician's Best Friend?
Imagine a politician is like a busy chef trying to run a massive kitchen (the government). They have thousands of recipes (laws) to manage, hungry customers (voters) to please, and only a few hours in the day. To keep up, they hire sous-chefs (advisors) to help them read recipes, suggest ingredients, and explain why a new dish should be served.
This paper asks: Can a super-smart AI (a Large Language Model) be that sous-chef?
The researchers tested this by asking six different AI models to explain why a new Romanian law should be passed. They then compared the AI's answers to the official, real-life explanations written by the actual Romanian government.
The Setup: The "Taste Test"
The researchers didn't just ask the AI "What do you think?" They gave it a specific challenge:
- The Task: Read a new law proposal.
- The Goal: Write a "reasoning memo" (an expuneri de motive) explaining why this law is good, just like a real government official would.
- The Judges: They used a mix of other AIs and computer programs to grade how close the AI's answer was to the real government's answer.
They tested six "chefs":
- The Top Chefs (Frontier Models): GPT-5 (Mini & Chat) and Claude Haiku 4.5. These are the most advanced, expensive, and powerful AIs.
- The Junior Chefs (Open-Weight Models): Three versions of Meta's Llama (small, medium, and huge). These are free to use but generally less powerful.
The Results: A Two-Tier Kitchen
The results were surprisingly clear. The AI chefs fell into two distinct groups, like a First-Class seat vs. Economy seat on a plane.
1. The "First-Class" Tier (The Pros)
The top three models (GPT-5 and Claude) were almost indistinguishable from each other. They scored a 4.6 out of 5.
- What this means: They sounded very much like real politicians. They used the right legal jargon, structured their arguments correctly, and sounded convincing.
- The Surprise: The smallest, cheapest model from Anthropic (Claude Haiku) performed just as well as the massive, expensive GPT-5 models. It's like a compact sports car driving just as fast as a luxury limo on this specific track.
2. The "Economy" Tier (The Juniors)
The three Llama models scored significantly lower (around 3.7 to 4.0).
- What this means: They were okay, but they missed the mark. They sounded a bit more generic and less precise.
- The Lesson: Making the AI "bigger" (adding more brain power/parameters) didn't help much. A 70-billion-parameter Llama didn't perform much better than an 8-billion one. The gap wasn't about size; it was about who taught them (the training data).
The Hidden Danger: The "Confident Liar" (The Competence Trap)
Here is the most important warning from the paper.
The top-tier AIs were great at the structure of the argument. They knew how to write a political memo. But, they sometimes made up the facts.
- The Analogy: Imagine a sous-chef who writes a perfect recipe for a "Spicy Tuna Dish." The ingredients list looks professional, the steps are logical, and the presentation is beautiful. But, the chef accidentally listed "Tuna" when the law was actually about "Salmon," or invented a fake law number that doesn't exist.
- The Risk: Because the AI sounds so fluent and confident, the politician (the boss) might not notice the mistake. This is called the "Competence Trap." The AI is so good at sounding smart that it tricks the human into trusting facts that are actually made up.
When Does AI Work? (The "Template" vs. The "Wild Card")
The researchers found that AI is a great helper for boring, predictable tasks, but a terrible helper for weird, political tasks.
- The Easy Mode (Templates): If a law is just copying a rule from the European Union (like "We must change the speed limit to match EU rules"), the AI is fantastic. It knows the template, fills in the blanks, and gets it right.
- The Hard Mode (Political Wild Cards): If a law is about something weird and specific, like "We need to give money to a specific local church because the Mayor likes them," the AI fails. It doesn't know the local gossip or the hidden political reasons. Instead, it makes up a generic, plausible-sounding reason that is actually wrong.
The Big Takeaway: It's Not About "Bias," It's About "Ignorance"
Most people worry that AI is "biased" (e.g., "It's too liberal" or "It's too conservative"). This paper says that's not the main problem for politicians.
The real problem is "Contextual Ignorance."
The AI isn't trying to trick you with a political agenda; it just doesn't know the local context. It's like a tourist trying to give you directions in a city they've never visited. They might use the right words ("Turn left at the big tree"), but if there is no big tree, they are just guessing.
Summary for the Everyday Person
- AI is getting scary good at sounding like a politician. The top models are almost as good as the real thing for standard tasks.
- Don't trust the facts blindly. Even the best AI can confidently invent fake laws or numbers. You still need a human to double-check the details.
- Size doesn't matter as much as you think. A tiny, cheap AI can sometimes do the job as well as a giant, expensive one, provided it was trained on the right data.
- The "Cascading Failure": The politician is busy (can't check everything), the AI is smart but ignorant (makes up facts), and the tools we use to check the AI are also imperfect. This creates a chain reaction where mistakes can slip through easily.
The Bottom Line: AI can be a great "drafting assistant" for politicians, but it should never be the "final decision-maker." It's a powerful tool, but like any tool, it needs a skilled human to hold the handle and watch where it cuts.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.