← Latest papers
⚡ electrical engineering

SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling

The paper introduces Sleeping-DISCO 9M, a novel large-scale open-source dataset composed of actual popular music and world-renowned artists designed to overcome the limitations of existing synthetic or non-representative datasets for generative music modeling.

Original authors: Tawsif Ahmed, Andrej Radonjic, Gollam Rabby

Published 2026-06-24
📖 4 min read☕ Coffee break read

Original authors: Tawsif Ahmed, Andrej Radonjic, Gollam Rabby

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to teach a robot to sing like a human. To do this, you need to feed it a massive library of songs, lyrics, and stories about music. For a long time, the "big tech" companies (like the ones behind Jukebox) kept their secret libraries locked in a vault, while smaller researchers had to make do with tiny, homemade collections or fake music.

Sleeping-DISCO 9M is a new, massive library that the authors have built and opened up for everyone to use. Here is the simple breakdown of what they did and why it matters, using some everyday analogies:

1. The Problem: The "Empty Pantry" vs. The "Secret Vault"

Think of the world of AI music research like a group of chefs trying to invent new recipes.

  • The Big Chefs (Big Tech): They have secret pantries full of the world's most famous ingredients (popular songs by famous artists), but they won't let anyone else see them.
  • The Small Chefs (Researchers): They have been trying to cook with tiny, artificial ingredients (synthetic music) or very small, specific collections (like only Chinese pop songs).
  • The "Random" Libraries: There were other big libraries available (like DISCO-10M), but they were like a warehouse full of unlabelled boxes. You knew there was music in there, but you didn't know who sang it, what the song was about, or even if the music was real or just a random clip. They were too messy to be useful for serious cooking.

2. The Solution: A "Super-Organized" Music Encyclopedia

The authors created Sleeping-DISCO 9M. Think of this not just as a pile of MP3 files, but as a massive, perfectly organized music encyclopedia.

  • The Scale: It contains nearly 9 million songs from over 648,000 different artists. That's like having a library that holds almost every hit song from the last decade.
  • The Quality: Unlike the messy warehouses mentioned above, this library is organized. Every song comes with a detailed "ID card" (metadata). You can search for a song by the artist, the album, the year it was released, or even the specific genre.
  • The Diversity: It's not just English pop. It's a global mix, covering 169 languages (from English and Japanese to Swahili and Hausa). It's like a world tour in a single dataset.

3. How They Built It: The "Digital Librarian"

The team didn't just download random files. They built a digital robot librarian (a Python scraper) that visited the Genius website (a famous site for lyrics and music facts).

  • The robot carefully read through millions of pages, copying down song titles, artist names, album details, and lyrics.
  • It then went hunting on YouTube to find the actual video links for these songs, using a "smart match" system to ensure the video title actually matched the song they were looking for.
  • The Catch: They couldn't share the actual lyrics or the deep "annotations" (the fun facts written by Genius editors) because those are copyrighted. However, they did share the links to the songs and the "fingerprint" (embeddings) of the lyrics, which helps computers understand the meaning without needing the raw text.

4. Why This is a Big Deal

Before this, if a researcher wanted to train an AI to understand real-world music, they had to build their own messy collection or use a tiny, limited one.

  • Sleeping-DISCO is the first time a dataset of this size and quality (featuring real, famous artists like Maroon 5 and Shakira) has been made open-source.
  • It bridges the gap between the "secret vaults" of big companies and the "small kitchens" of independent researchers.

5. The Rules of the Library

The authors are sharing this library under a specific set of rules (CC-BY-NC-ND 4.0).

  • You can use it to study and build your own music models.
  • You cannot sell it or make money from the dataset itself.
  • You cannot change the dataset and claim it's your own.
  • The Lyrics: The actual text of the songs is still locked up (because of copyright), but the authors will share it with universities and researchers who promise to use it strictly for science.

In short: The authors built a giant, clean, and diverse music library that researchers can finally use to teach AI how to understand and generate real music, filling a gap that has existed for years.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →