← Latest papers
🤖 machine learning

Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

This study demonstrates that general-purpose pretrained audio foundation models inherently encode deep phylogenetic relationships in marine mammal and bird vocalizations, outperforming both hand-crafted features and domain-specific models without requiring taxon-specific pretraining.

Original authors: Víctor Rincón Yepes

Published 2026-07-27
📖 6 min read🧠 Deep dive

Original authors: Víctor Rincón Yepes

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Sound of Family Trees

Imagine you are walking through a giant, noisy forest where every animal has a unique voice. Some chirp, some roar, and some sing complex songs. For a long time, scientists have wondered: if you listen closely, can you tell who is related to whom just by how they sound? This question sits at the intersection of two big ideas in science. The first is evolution, the story of how species change and split apart over millions of years, like branches growing on a giant tree. The second is bioacoustics, the study of animal sounds.

The big mystery is whether the "family tree" of an animal is written into its voice. If two species sound very similar, does that mean they are close cousins? Or could they just be neighbors who learned to shout the same way? To answer this, scientists used to measure sounds with simple tools, looking at basic features like pitch or the "texture" of the noise. But recently, computers have gotten much smarter. We now have "foundation models"—massive AI programs trained on millions of hours of human speech, music, and nature sounds. These AIs can understand sound in a very deep, complex way, almost like a human listener. The big question this paper asks is: Do these super-smart computers, which were never taught about animal families, accidentally learn to recognize evolutionary relationships just by listening to the sounds?


The Study: Listening to the Family Tree

In this study, researchers decided to test if these modern AI "ears" could hear the hidden history of animal families. They looked at two very different groups of animals: 32 species of marine mammals (like whales, dolphins, and seals) and 20 species of birds.

Think of the AI models as different types of detectives. The researchers gave them a massive library of recordings and asked them to measure how "similar" the sounds of different species were. They compared these sound-similarity scores against the actual "family tree" distances (how many millions of years ago two species shared a common ancestor).

They tested four different ways of analyzing the sound:

  1. The Old Way (MFCC): This is like using a basic ruler to measure sound. It looks at simple, hand-made features.
  2. The General AI (AST & CLAP): These are huge, general-purpose AI models trained on everything from car engines to pop songs. They weren't taught about animals at all.
  3. The Specialist AI (BEATs-bio & BirdNET): These are models specifically trained on animal sounds or bird calls, hoping they would be better at spotting family connections.

The Big Surprise: The Generalists Won

The results were a bit of a shock. The "Old Way" (the basic ruler) barely found any connection between how animals sound and how they are related. It was like trying to read a book by looking at the color of the paper; it just didn't work.

However, the General AI models (the ones trained on human music and speech) did something amazing. They found a strong signal. When the researchers looked at whales and dolphins, the AI could tell that species that split apart more recently sounded more alike, and those that split apart long ago sounded very different. In fact, for the whale family, the connection was incredibly strong.

Here is the twist: The Specialist AIs (the ones trained specifically on birds or bioacoustics) did not do better than the general ones. In fact, for the birds, the general-purpose models were actually slightly better at finding the family tree than the bird-expert models. The paper suggests that training a model specifically on animals might actually make it worse at seeing the big evolutionary picture, perhaps because it gets too focused on tiny details or specific species labels rather than the deeper structure of the sound.

What the AI Actually Heard

The researchers wanted to make sure the AI wasn't just cheating by noticing simple things, like "big animals have low voices and small animals have high voices." They ran special tests to see if the AI was just listening to the pitch. The answer was no. Even when they removed the pitch from the equation, the AI still found the family connections.

This means the AI is picking up on something much more subtle: the style of the call. It's like how you can tell a family member is talking even if they are whispering or shouting, because of how they shape their words, not just how loud they are. The AI learned the "accent" of evolution.

The Mix-Up: When Sound Lies

The study also found some funny exceptions that prove the rule.

  • The Sound-Alike Strangers: The Bearded Seal and the Northern Right Whale are cousins who split apart 90 million years ago. They should sound totally different. But they sound almost identical! The AI noticed they sound the same, even though they are ancient strangers. This is likely because they both live in similar cold environments and need to make low, rumbling sounds to travel through water.
  • The Sound-Different Cousins: The Leopard Seal and the Weddell Seal are very close cousins, splitting apart only 10 million years ago. You'd expect them to sound similar. But they sound completely different! One makes complex, musical underwater songs, while the other makes low growls.

These examples show that while evolution leaves a mark on sound, the environment can sometimes force animals to change their voices so much that they sound like strangers, or so similar that they sound like family.

The Bottom Line

This paper tells us that we don't need to build special, animal-only computers to understand the evolutionary history of animal voices. The general-purpose AI models, trained on the sounds of our human world, are already incredibly good at hearing the "family tree" hidden in animal calls. They are better than the old, simple tools, and surprisingly, they are even better than the models built specifically for animals.

The researchers suggest that the secret isn't in training the AI on animals, but in the way these general models are built to understand the complex patterns of sound. They propose that future experiments should test if this is because of the AI's brain structure, the way it learns, or the huge variety of sounds it was trained on. But for now, the message is clear: if you want to listen to the history of life in a sound, you might just need to ask a general-purpose AI to listen for you.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →