PREMOVE: A multilayer manually annotated dataset of PREverbed MOtion VErbs in Ancient Greek and Latin
This paper introduces PREMOVE, a publicly available, multilayer manually annotated dataset of preverbed motion verbs in Ancient Greek and Latin, designed to support cross-linguistic, diachronic, and computational research through its comprehensive morphosyntactic and semantic annotations of 2,835 verbal occurrences across 35 texts.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery about how ancient people talked about movement. Specifically, you want to understand how they used "prefixes" (little word bits added to the front of verbs, like under- in undergo or re- in return) to change the meaning of words like "go," "run," or "sail."
The paper introduces PREMOVE, a massive, meticulously organized digital filing cabinet designed to help researchers solve this mystery in two ancient languages: Ancient Greek and Latin.
Here is a breakdown of what the paper claims, using simple analogies:
1. The Problem: A Scattered Library
Before this project, researchers had many books (corpora) of ancient Greek and Latin texts. However, these books were like libraries where the books were sorted by author or date, but not by the specific "action" of the words inside. If you wanted to find every instance of "going away" or "sailing around" across thousands of years of history, you'd have to read every single page manually. There was no single map to show how these "prefixes" changed the meaning of motion verbs across different times and genres.
2. The Solution: The PREMOVE Filing Cabinet
The author, Andrea Farina, built PREMOVE. Think of this as a giant, 42-layer spreadsheet containing 2,835 specific examples of motion verbs that have been tagged with a prefix.
- The Ingredients: The data comes from a carefully selected "Base Corpus" of 35 texts (like the Iliad, Caesar's war journals, or Cicero's speeches). The author balanced these texts to ensure they represent different time periods (from the 8th century BCE to the 2nd century CE) and different styles (poetry, history, drama, philosophy).
- The Ingredients List: They picked 8 common motion verbs per language (like "go," "flee," "run," "sail") and 16 common prefixes (like "out," "in," "around," "under").
- The Layers: Every single example in the spreadsheet is annotated with 42 different pieces of information. Imagine a photo of a car accident. A basic photo shows the car. PREMOVE is like a photo where every part is labeled: the speed, the weather, the road type, the driver's mood, and the exact location on a map.
- Morphology: What form is the word in?
- Semantics: Does it mean "going up" or "going down"? Is it literal or metaphorical?
- Spatial Roles: Is the word describing a starting point (Source), a destination (Goal), or a path?
- Context: Who wrote it? When? What genre?
3. The Quality Control: The "Three Judges" Test
Because humans are annotating this data, there's always a risk of disagreement. To prove the data is reliable, the author tested it.
- The Test: Three different experts (a computer linguist, a traditional historian, and the author) looked at a small sample of sentences and tried to label them independently.
- The Result: For the most important parts—identifying the prefix and its basic meaning—they agreed almost perfectly (like three judges all giving a contestant a 10/10). For more complex, subjective meanings (like the exact emotional nuance of a verb), the agreement was lower, which is normal for human interpretation. This proves the system works well for its core purpose.
4. The Tools: A Map and a Telescope
The paper describes two main ways to use this data:
- The Dataset (CSV): This is the raw data, available for download. It's like giving a researcher a box of raw ingredients so they can cook whatever they want.
- PrevNet (The Interface): This is a user-friendly website built on top of the data. It's like a "Google Maps" for ancient language. You can click on a prefix (like "around") and instantly see a pie chart showing how often it was used, which authors used it, and how its meaning shifted over centuries. You can click a slice of the pie and see the actual ancient sentences.
- The Connection (LiLa): The data is also linked to a massive global network of language data (LiLa), making it "FAIR" (Findable, Accessible, Interoperable, Reusable). This is like giving the data a universal barcode so any computer system can read it.
5. What It's Used For (According to the Paper)
The paper explicitly states this dataset is a tool for:
- Comparing Languages: Seeing how Greek and Latin handled movement differently.
- Training AI: Using the data to teach Large Language Models (LLMs) how to understand ancient languages better.
- Historical Linguistics: Tracking how words changed meaning over 1,000 years (e.g., how a word for "going out" might eventually mean "happening").
- Digital Humanities: Helping students and scholars explore texts without needing to be coding experts.
What It Is NOT (Based on the Paper)
- It is not a medical or clinical tool.
- It does not claim to solve modern language learning problems directly (though it helps train the AI that might).
- It is not a complete dictionary of every word in Greek or Latin; it is a focused study on motion verbs with prefixes.
In short, PREMOVE is a high-precision, human-verified map of how ancient people moved through the world using words, designed to help computers and humans understand the history of language together.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.