← Latest papers
⚡ electrical engineering

CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning

This paper introduces CultureMERT-95M, a multi-culturally adapted music foundation model trained via a novel two-stage continual pre-training strategy on diverse non-Western traditions, which significantly improves cross-cultural auto-tagging performance while maintaining proficiency on Western benchmarks and offering a viable alternative through task arithmetic.

Original authors: Angelos-Nikolaos Kanatas, Charilaos Papaioannou, Alexandros Potamianos

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Angelos-Nikolaos Kanatas, Charilaos Papaioannou, Alexandros Potamianos

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Music is a universal language, yet the computers we use to understand it often speak only one dialect. For years, the most powerful artificial intelligence models designed to listen to and categorize music have been trained almost exclusively on Western styles, such as pop, rock, and classical symphonies. These systems have become incredibly good at recognizing the patterns of those specific traditions, but they struggle when faced with the vast, rich diversity of the world's other musical cultures. They miss the subtle melodic structures of Indian classical music, the unique rhythmic cycles of Turkish traditions, or the distinct scales of Greek folk songs. This limitation means that for billions of people whose musical heritage falls outside the Western canon, these digital tools remain blind, unable to preserve, recommend, or analyze the art that defines their communities.

Researchers have long sought a way to teach these models to listen more broadly without having to rebuild them from the ground up, a process that would be prohibitively expensive and time-consuming. The challenge lies in adapting a model that has already learned one set of rules to understand a completely different set of rules without causing it to forget what it already knew. This is the central problem addressed by a new study introducing a system called CultureMERT. The researchers aimed to take a foundation model that had been trained on a thousand hours of predominantly Western music and gently guide it to understand the musical languages of the Eastern Mediterranean and the Indian subcontinent, all while ensuring it did not lose its ability to recognize the Western styles it was originally built to know.

To achieve this, the team employed a strategy known as continual pre-training. Instead of discarding the existing model and starting over, they fed it a new, carefully balanced diet of audio data. This new dataset included 650 hours of music drawn from four distinct traditions: Turkish makam, Hindustani classical, Carnatic classical, and Greek traditional music. However, simply feeding this new data to the model caused it to stumble, a phenomenon where the computer's performance on its original tasks would crash as it tried to learn the new ones. To solve this, the researchers devised a two-stage approach. In the first stage, they introduced the model to a smaller portion of the new music while keeping its most complex internal layers frozen, allowing only the basic sound-processing parts to adjust. They also mixed in a small amount of Western music during this phase to remind the model of its original training. In the second stage, they unlocked the entire model and trained it on the full 650-hour dataset, allowing it to fully integrate these new cultural nuances.

The results were striking. When tested on tasks that required the model to identify instruments, moods, or specific musical modes in non-Western music, the adapted model showed a significant improvement, averaging a nearly five percent increase in accuracy compared to the original version. Crucially, this gain did not come at the cost of its previous knowledge; the model retained its ability to recognize Western music with almost no loss in performance. The study also explored a different method called task arithmetic, which involves mathematically combining the "knowledge" of several smaller models, each trained on a single culture, into one unified system. This approach performed just as well as the direct training method on non-Western tasks and even outperformed it on Western benchmarks, suggesting that merging specialized models is a powerful alternative to retraining them all at once.

The researchers found that the model's ability to transfer knowledge between cultures was not uniform. For instance, the model trained on Carnatic music from South India showed a strong ability to generalize to Hindustani music from North India and even to Turkish traditions, likely because these systems share deep theoretical roots in melody and rhythm. In contrast, the model struggled more to transfer knowledge to Greek traditional music, which, while distinct, shared fewer structural similarities with the Indian and Turkish datasets in the specific way the computer encoded sound. By analyzing the digital "tokens" the model uses to represent sound, the team discovered that the mathematical distance between these musical traditions predicted how well the model would learn them, offering a new way to plan future training data.

Ultimately, this work demonstrates that artificial intelligence can be made more inclusive without sacrificing its core capabilities. The researchers have released their adapted model, CultureMERT, to the public, providing a tool that can finally listen to the world's music with a much broader ear. They acknowledge that while their model is a significant step forward, it is not a perfect solution; the underlying technology used to convert sound into digital tokens was still trained on Western music, which may limit how deeply it can grasp certain cultural subtleties. Nevertheless, by showing that a model can learn new musical languages while remembering the old ones, this study offers a clear path toward a future where digital music systems are truly global, capable of honoring the full spectrum of human musical expression.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →