← Latest papers
💻 computer science

A global dataset of continuous urban dashcam driving

The paper introduces CROWD, a globally diverse, manually curated dataset of over 20,000 hours of continuous, unedited urban dashcam footage from 238 countries, designed to support robust cross-domain driving research by providing routine driving segments with manual time-of-day and vehicle labels alongside machine-generated object detection and tracking annotations.

Original authors: Md Shadab Alam, Olena Bazilinska, Pavlo Bazilinskyy

Published 2026-04-08
📖 4 min read☕ Coffee break read

Original authors: Md Shadab Alam, Olena Bazilinska, Pavlo Bazilinskyy

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to teach a robot how to drive a car. To do this, you need to show it millions of hours of driving footage so it can learn what a normal street looks like, how people walk, and how other cars behave.

Most existing driving datasets are like highlight reels from a sports channel. They only show the exciting, dangerous, or weird moments: car crashes, near-misses, police chases, and traffic jams. While these are dramatic, they don't represent the boring, everyday reality of driving 99% of the time. If you only train a robot on crash videos, it might think the world is a constant battlefield!

Other datasets are like high-end, expensive spy missions. They use special cars with lasers and cameras mounted all over them, driving only in a few specific cities (like San Francisco or Munich). These are great, but they are expensive to make, cover very little ground, and don't look like the view from a regular person's dashboard.

Enter CROWD: The "Everyday Driving" Library

The paper introduces CROWD (City Road Observations With Dashcams). Think of CROWD as a massive, global library of boring, normal, everyday driving videos collected from YouTube.

Here is how they built it, using simple analogies:

1. The Source: Finding the "ASMR" of Driving
Instead of looking for crash videos, the researchers went hunting for a specific type of YouTube video: long, continuous, unedited dashcam footage. They looked for videos that people upload just to relax to (often called "ASMR" driving videos). These are the videos where nothing exciting happens; the car just drives through a city for an hour. This is exactly the "routine" data needed to teach a robot what normal driving feels like.

2. The Filter: The "Safety Editor"
The team manually watched thousands of these videos and acted like strict editors. They cut out anything that wasn't "normal":

  • No Crashes: If a car hit something, that clip was tossed.
  • No Editing: If the video had jump cuts or fancy transitions, it was discarded. They wanted continuous, real-time flow.
  • No Non-City Areas: They only kept videos driving through towns and cities, not on empty highways or in the middle of a cornfield.
  • No Stopped Cars: If the car was parked at a gas station for 10 minutes, that part was cut. They only wanted the car moving in traffic.

3. The Result: A Global Mosaic
The final dataset is huge. It contains over 20,000 hours of driving footage from 7,103 different cities across 238 countries.

  • Analogy: Imagine taking a snapshot of a street in Tokyo, then one in Nairobi, then one in Buenos Aires, and stitching them all together into one giant, continuous movie. That's CROWD. It covers almost every corner of the globe, from the Americas to Africa to Asia.

4. The "Smart Glasses" (Annotations)
Just giving someone a raw video isn't enough for a computer to learn. The computer needs to know what it is looking at.

  • The researchers ran a super-smart AI (called YOLO) over every single frame of these videos.
  • This AI put "stickers" (bounding boxes) around everything: people, bicycles, cars, buses, traffic lights, and stop signs.
  • It also tracked them, meaning it knows that "that red car" at second 10 is the same "red car" at second 20.
  • They didn't give away the actual video files (to respect copyright and privacy); instead, they gave researchers the "map" (timestamps and IDs) so anyone can download the original YouTube video and use the "stickers" to study it.

Why Does This Matter?

The "Generalist" Driver
Most AI models are trained on data from just one or two cities. If you take a model trained in sunny California and drive it in rainy London, it might get confused. CROWD is like a universal driver's ed course. Because it includes so many different countries, weather conditions, and road styles, it helps train AI that can handle any city in the world, not just the ones where the data was collected.

The "Safety" Focus
By removing all the crashes and focusing on routine driving, CROWD helps researchers answer a different question: "How often do cars and pedestrians actually interact in a normal day?" This helps us understand risk before accidents happen, rather than just studying them after they occur.

In Summary

CROWD is a global, open-source library of boring, normal driving videos from YouTube, cleaned up and labeled by computers. It fills the gap between "crash highlight reels" and "expensive, limited city datasets," providing a realistic, diverse foundation for teaching autonomous vehicles how to navigate the real, messy, everyday world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →