Counting trees: A treebank-driven exploration of syntactic variation in speech and writing across languages
This paper introduces a novel, inductive treebank-driven framework using Universal Dependencies to demonstrate that spoken language across English and Slovenian relies on a smaller, less diverse, and largely non-overlapping set of syntactic structures compared to writing, reflecting modality-specific demands for interactivity and economy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to understand how people build things. In this case, the "things" are sentences, and the "building blocks" are words and the invisible rules that hold them together.
For a long time, linguists (language detectives) mostly looked at written sentences. They treated speech as just a messy, informal version of writing. But this paper, titled "Counting Trees," argues that speech and writing are actually two different species of the same animal, built with different blueprints.
Here is the simple breakdown of what the researchers did and what they found, using some everyday analogies.
1. The New Tool: The "Tree" Scanner
Traditionally, to study grammar, linguists would look for specific patterns, like "subject-verb-object." It's like looking for only red bricks in a building.
This paper introduces a new, super-powerful scanner. Instead of looking for specific bricks, they treat every sentence as a family tree (which is why they call it a "treebank").
- The Tree: Every word in a sentence is a branch. The main verb is the trunk. The other words are branches and leaves.
- The Method: They didn't just look at the whole tree. They chopped every single sentence into every possible smaller tree inside it. They took a sentence like "The cat sat on the mat," and they extracted the tree for "cat," the tree for "sat," the tree for "on the mat," and the whole thing combined.
- The Goal: They did this for both Speech (people talking) and Writing (books, news, essays) in two very different languages: English and Slovenian.
2. The Big Discovery: The "Fast Food" vs. "Fine Dining" Kitchen
When they compared the "trees" found in speech versus writing, they found a massive difference in variety.
- Writing is Fine Dining: The written language had a huge, diverse menu. It used a vast, complex, and varied set of tree structures. It was like a chef using hundreds of different techniques to create unique, elaborate dishes.
- Speech is Fast Food: The spoken language was much more repetitive. It relied on a smaller, simpler set of "trees" that were used over and over again. It was like a fast-food kitchen using the same few assembly lines to make thousands of burgers quickly.
The Analogy: Imagine a library.
- Writing is like a library with 10,000 unique, complex books.
- Speech is like a library where 90% of the books are just copies of the same 500 popular titles, with maybe a few new ones added every day.
3. The Shocking Result: They Don't Even Share Recipes
The researchers expected speech and writing to share some common structures. They thought, "Surely, we use the same basic sentence structures, just with different words."
They were wrong.
They found that the overlap between speech and writing was tiny.
- The "Ghost" Structures: Most of the tree structures found in spoken language did not exist in the written language at all.
- The Analogy: It's like if you asked a chef to write down a recipe for a "hamburger," and they wrote a recipe for a "steak." They are both food, but the instructions (the syntax) are completely different. Speech uses "recipes" (sentence structures) that are so specific to talking that you simply don't see them in books.
4. Why Does This Happen? The "Real-Time" Pressure
Why is speech so different? The paper suggests it's because of pressure.
- Writing: You have time to think, edit, and build complex, multi-layered trees. You can plan a sentence like an architect planning a skyscraper.
- Speech: You have to build the sentence while you are speaking. You can't stop to edit. So, you use "prefabricated" structures—short, simple, repetitive trees that are easy to assemble on the fly.
- Example: Instead of saying, "I am going to the store because I need milk," a speaker might say, "Going to store. Need milk." The "tree" is shorter and simpler because the speaker is building it in real-time.
5. The "Language Independence" Trick
The researchers tested this on English (which relies heavily on word order) and Slovenian (which uses word endings and can scramble word order).
Even though these two languages are totally different, the pattern was the same:
- In both languages, speech was simpler and more repetitive than writing.
- In both languages, speech used unique structures that didn't appear in writing.
This proves that the difference isn't just about the language itself; it's about the mode (speaking vs. writing). It's a universal rule of human communication.
Summary: What Does This Mean for Us?
This paper is a game-changer because it stops us from treating speech as just "bad writing."
- Old View: Speech is just writing with more mistakes and pauses.
- New View: Speech is a completely different system with its own unique grammar, its own "trees," and its own rules.
The Takeaway:
If you want to understand how humans really communicate, you can't just read books. You have to listen to how people talk. The "trees" they grow in conversation are shorter, simpler, and more unique than the trees they grow on the page. This new method gives us a map to finally see those hidden trees clearly.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.