← Latest papers
💻 computer science

Teach a Molmo2Fish: Towards interactive fish tracking with natural language guidance

This paper introduces Molmo2Fish, an interactive workflow that leverages a multimodal large language model to enable natural language-guided correction and tracking of fish in sonar datasets, demonstrating high performance while highlighting opportunities for further improvement in language integration.

Original authors: Kai Van Brunt (Massachusetts Institute of Technology), Justin Kay (Massachusetts Institute of Technology), Sara Beery (Massachusetts Institute of Technology)

Published 2026-08-20✓ Author reviewed
📖 5 min read🧠 Deep dive

Original authors: Kai Van Brunt (Massachusetts Institute of Technology), Justin Kay (Massachusetts Institute of Technology), Sara Beery (Massachusetts Institute of Technology)

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the quiet, dark waters of rivers where salmon return to spawn, scientists face a difficult counting problem. To protect these fish populations and set sustainable fishing limits, researchers need a reliable estimate of how many fish pass a specific point each year. They use underwater sonar cameras to record this movement, but the resulting videos are often murky, filled with noise, and crowded with hundreds of fish swimming at once. For years, people have had to review this footage manually to count the fish, a slow and tedious process. Recently, computers have learned to do this automatically, but they often stumble when the water conditions change or when too many fish appear at once, losing track of individuals or counting the same fish twice. This is where a new kind of artificial intelligence comes in, one that does not just try to be perfect on its own, but learns to listen to a human voice to fix its mistakes.

A team of researchers at the Massachusetts Institute of Technology has developed a tool called Molmo2Fish to bridge the gap between imperfect computer vision and the needs of ecologists. Instead of building a system that tries to get everything right the first time, they created a conversational partner. This system is based on a multimodal large language model, a type of artificial intelligence that can see images and understand text at the same time. The researchers taught this model to watch sonar videos of salmon and predict where each fish is moving. But the true innovation lies in what happens next: if the computer misses a fish or loses track of one, a human can simply type a sentence like, "You are missing a fish near the bottom that is swimming left," and the model will immediately adjust its predictions to correct the error. It is a shift from a rigid, automated process to a flexible conversation where the computer and the human work together to build an accurate record.

To test this idea, the team took the Molmo2 model, which was originally trained on a visually diverse set of videos including people, animals, and automobiles, and specialized it for the unique, grainy world of underwater sonar. They fed it clips from the Caltech Fish Counting dataset, the largest open collection of fisheries sonar video, teaching it not only how to track fish but also how to accept corrections. The researchers created a training process where the model would make a first guess at the fish paths, and then a simulated human would point out the errors in plain English. The model learned to listen to these instructions and update its mental map of the video in real time. They found that the system could successfully learn this new skill, transforming from a tool that struggled with sonar data into one that could track fish with high accuracy, especially when given a little help.

The results showed that this interactive approach works well, particularly when the initial computer predictions are far off. When the model started with a basic, traditional tracking system that often missed fish or broke a single fish's path into several confusing fragments, Molmo2Fish could fix those errors effectively after a single round of human feedback. In some river locations where the fish were small and hard to see, the corrected tracks were significantly better than what the computer could achieve alone. However, the researchers also discovered a limit to this power. When the computer was already doing a very good job on its own, asking it to correct its own work with more instructions did not always make it better. In fact, sometimes adding more specific instructions confused the model, leading it to make changes that were unnecessary or even wrong.

The study also revealed that the model is highly sensitive to the type of feedback it receives. It learned to pay close attention to specific details, such as which fish to merge or which ones to ignore, but it struggled when the instructions were vague or when the errors were of a type it had never seen before. For instance, if the model was trained only on certain kinds of mistakes, it failed to recognize new, different types of errors when they appeared in the test videos. This indicates that while the technology is promising, it still needs more diverse training data to become truly robust. The researchers suggested that this method of "teaching a model to fish" is a promising direction for handling the messy reality of ecological data, with the goal of eventually allowing non-experts to get high-quality results without needing to retrain complex computer systems from scratch, though the method is not yet reliable enough to deliver that on a typical video.

Ultimately, Molmo2Fish represents a step toward a future where artificial intelligence in science is not a black box that produces a final answer, but a collaborative tool that evolves through conversation. By allowing researchers to guide the model with natural language, the system becomes adaptable to the specific challenges of each river and each season. The work demonstrates that we do not always need a perfect algorithm to solve a difficult problem; sometimes, a system that knows how to listen and correct itself is far more useful. As the researchers continue to refine this approach, the hope is that these tools will become standard in conservation, helping to protect salmon populations by turning hours of difficult counting into a manageable, interactive task.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →