← Latest papers
📄 other

A parallel treebank of learner Swedish

This paper introduces a parallel treebank of Swedish learner language based on the SweLL corpus, annotated according to the Universal Dependencies standard to address L2-specific challenges, evaluate annotation quality, and identify factors influencing annotation difficulty.

Original authors: Arianna Masciolini, Maria Irena Szawerna, Aleksandrs Berdicevskis, Caroline Grand-Clement, Elena Volodina

Published 2026-07-01
📖 5 min read🧠 Deep dive

Original authors: Arianna Masciolini, Maria Irena Szawerna, Aleksandrs Berdicevskis, Caroline Grand-Clement, Elena Volodina

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Building a "Bilingual Dictionary" for Mistakes

Imagine you are trying to learn a new language, like Swedish. You write a sentence, but it's full of mistakes. A teacher reads it, fixes the errors, and writes down the "correct" version. Now, imagine you have a massive library of these "Before" (your messy draft) and "After" (the teacher's clean version) pairs.

This paper introduces a new digital library called UD Swedish-SweLL. It is a collection of 510 pairs of sentences written by adults learning Swedish. The authors didn't just store the text; they built a detailed "map" for every sentence, showing how the words connect to each other. They call this a treebank.

Think of a treebank like a family tree for a sentence. It shows who is the "parent" (the main word) and who are the "children" (the words that depend on it). For example, in "The big dog," "dog" is the parent, and "big" is a child attached to it.

Why Do This? (The "Universal Translator" Idea)

Usually, when researchers study learner mistakes, they just make a list of errors: "Spelling mistake," "Wrong verb," "Bad grammar." It's like a mechanic saying, "The car is broken," without explaining how the engine works.

The authors decided to use a system called Universal Dependencies (UD).

  • The Analogy: Imagine UD is like a universal remote control that works for every TV brand (every language). Instead of inventing a new remote for Swedish, they use the same one used for English, Spanish, and Chinese.
  • The Benefit: Because they use this standard "remote," they can easily compare Swedish learners with learners of other languages. They can also use computer programs (parsers) that already know how to use this remote to help analyze the data faster.

The Challenge: Mapping a "Broken" House

The tricky part is that learner sentences are often "broken." They might have words spelled wrong, missing pieces, or words stuck together that shouldn't be.

  • The Problem: If you try to map a sentence that says "I go store" instead of "I go to the store," a standard map might get confused.
  • The Solution: The authors created a special set of rules (guidelines) for this specific project.
    • Rule 1: Don't fix the map, fix the description. If a student writes "traffik" (misspelled traffic), the map keeps the spelling "traffik" but labels it as a noun. They don't pretend it was spelled correctly; they document the mistake as it happened.
    • Rule 2: Follow the intent. If a student writes a sentence that looks like a direct translation from their native language (like Arabic or Kurdish), the authors analyze the Swedish words based on how the student intended them to work, even if the grammar is weird.

The Process: Humans and Robots Working Together

The team didn't just let a computer do all the work. They used a "Human-in-the-Loop" approach:

  1. The Robot (UDPipe): First, a computer program took a quick look at the sentences and drew a rough draft of the maps. It was like a student sketching a rough outline of a drawing.
  2. The Humans (The Editors): Then, human experts (who are also learning Swedish!) went through the sketches. They checked the computer's work, fixed the mistakes, and made sure the "Before" and "After" sentences matched up perfectly.
    • Analogy: Imagine the computer is a GPS that gives you a route. The human is the driver who knows that the GPS missed a pothole or a detour, so they correct the route before you drive.

Did It Work? (The Report Card)

The authors tested how well the humans agreed with each other and how well the computer did compared to the humans.

  • Human Agreement: The humans agreed with each other 98% to 99% of the time. This is like having three judges in a competition who almost always give the same score. It proves their rules were clear and easy to follow.
  • Computer vs. Human: The computer was pretty good, especially at identifying parts of speech (like knowing a word is a noun or a verb). However, when it came to figuring out the complex relationships between words (the "tree" structure), the computer made more mistakes than the humans.
  • The "Error Density" Factor: The computer struggled the most when the learner's sentence was full of errors. The more mistakes in the sentence, the harder it was for the computer to draw the map. Humans, however, were much better at handling messy sentences.

What's Next?

The authors say this is just the beginning. They have 510 sentences now, but the original library has over 1,000 essays. They plan to:

  1. Expand the treebank to include more sentences.
  2. Train the computer to get better by teaching it on these new, corrected maps.
  3. Eventually, use this system to automatically detect and analyze learner errors without needing a human to check every single sentence.

In short: They built a high-quality, standardized map of Swedish learner sentences. They proved that humans can agree on how to map these messy sentences, and they showed that while computers are getting good at it, they still need human help when the sentences get really tricky.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →