AI Super Simplified
Edition 309

The AI That “Talks to Dolphins” Can’t Translate a Word — and Still Isn’t Released | Edition 309

Edition 309 — DolphinGemma finds patterns in dolphin sound. It doesn’t translate, and 15 months after launch it’s still in development.

By Jerry Croteau

In April 2025, Google announced DolphinGemma — an AI model built with Georgia Tech and the Wild Dolphin Project to study the sounds wild dolphins make. The coverage wrote itself. Newsweek ran “Google Launches AI That Talks to Dolphins.” Other outlets suggested a conversation with another species was now a matter of months.

Fifteen months later, here is where it actually stands. Google DeepMind’s own page for the model says DolphinGemma “is currently in development” and “will be openly available” — future tense, both times. There is no download. There is no peer-reviewed result. And the model was never built to translate anything in the first place.

That reads like a debunking. It isn’t. The honest version of this story is stranger and more useful than the headline version, and it turns on a distinction that reaches well past dolphins: predicting what comes next is not the same as understanding what it means.

What was actually announced

DolphinGemma was unveiled on 14 April 2025 — National Dolphin Day, which tells you something about how carefully the launch was staged. Three parties built it: Google, the Georgia Institute of Technology, and the Wild Dolphin Project.

The Wild Dolphin Project is the reason any of this is possible. Founded in 1985, it is the world’s longest-running underwater dolphin research project, following the same community of Atlantic spotted dolphins in the Bahamas across generations. Its 42nd summer field season opened on 2 July 2026. What it has accumulated is not just audio — it is decades of underwater recordings painstakingly matched to individual named dolphins, their life histories, and what each animal was doing at the moment it made a sound. Almost nothing else in animal-communication research looks like that.

The model itself is deliberately small: around 400 million parameters, which is tiny by 2026 standards and entirely the point. It is designed to run on a Pixel phone in the water, alongside the researchers — the field hardware has been built around a Pixel 6, with a Pixel 9 generation planned. It is audio-in, audio-out: it takes dolphin sound, compresses it into tokens using a technique called SoundStream, and predicts what is likely to follow.

What “predicting the next sound” actually buys you

Google’s own description is precise, and worth reading slowly. The model “processes sequences of natural dolphin sounds to identify patterns, structure and ultimately predict the likely subsequent sounds in a sequence” — “much like how large language models for human language predict the next word or token in a sentence.”

That is a real scientific instrument, and the reason is not obvious. The hardest problem in studying an unknown communication system is that you cannot tell where one unit stops and the next begins. Dolphin sound is a continuous stream of whistles, burst-pulse squawks and click trains. Before you can ask what any of it means, you have to work out what the pieces are. A model that gets good at predicting the next sound has, in the process, learned where the boundaries fall and which sequences recur — because you cannot predict a stream well without internalising its structure.

Here is the catch, and it is not a small one. A model can reach near-perfect prediction on a signal while having no idea what any part of it refers to. Try it yourself — it takes about a minute.

Ten rounds of guessing what comes next in a sequence you have never seen before — then one question you cannot answer. · Open full-screen ↗

The genuinely astonishing part predates the AI

Dolphins call each other by name. That is not a headline gloss — it is a specific, replicated, peer-reviewed finding.

Each bottlenose dolphin develops a unique “signature whistle” early in life. In 2013, Stephanie King and Vincent Janik at the University of St Andrews published work in PNAS showing that dolphins respond to copies of their own signature whistle, and not to copies of other dolphins’ whistles played back the same way. Related work by King, Laela Sayigh, Randall Wells, Wendi Fellner and Janik in Proceedings of the Royal Society B found dolphins copying one another’s signature whistles, apparently to address specific individuals.

Learned, individually distinctive vocal labels, used to address one another, in a mammal that is not us. No AI was involved in establishing any of it.

It is worth holding that against the sweeping claims. Arik Kershenbaum, a zoologist at Girton College, Cambridge who studies animal communication, put the limit bluntly to Scientific American: “It’s not immediately clear that dolphins have words.” Having names is not the same as having a vocabulary, and a vocabulary is not the same as a language. His broader point is that human language is generative and effectively infinite — a fixed label for every object in your environment would not qualify.

CHAT is a different machine — and it is the one that “talks”

Most of the confusion in the coverage comes from merging two separate systems.

DolphinGemma listens to natural dolphin sound and looks for structure. CHAT — Cetacean Hearing Augmentation Telemetry, built by Thad Starner’s team at Georgia Tech — does something close to the opposite. It does not decode dolphin communication at all. It invents a small shared vocabulary of synthetic whistles, deliberately unlike anything the dolphins already say, each tied to an object the animals enjoy: a scarf, a piece of sargassum seaweed. Researchers play the whistle while passing the object around. If a dolphin mimics it, the system recognises the mimic in real time and the researcher hands the object over.

That is not translation. It is an attempt to build a tiny pidgin from scratch, and the Wild Dolphin Project is admirably unsentimental about how far it has got. Describing the moment a dolphin mimicked the sargassum whistle, their own write-up says flatly that “this is not to say that a dolphin knew what it was saying,” and that there was “nothing to indicate that this ‘word’ was used in context.” As they put it: “to mimic a whistle is not the same as functionally understanding how a sound can be used to communicate.”

DolphinGemma’s practical job in the field is to support CHAT — spotting a mimic earlier in the vocalisation, so the researcher can respond fast enough for the dolphin to connect the sound with what happens next.

Thea Taylor, managing director of the Sussex Dolphin Project, who is not involved in the work, flagged the obvious risk to Scientific American: researchers have to be careful they are not simply training the dolphins. The difference between an animal using a word and an animal repeating a sound because a reward follows is the entire question — and it is an easy one to fool yourself about.

What the coverage saidWhat the record shows
AI can now translate dolphin languageGoogle describes the model as identifying patterns and predicting the next sound. No translation capability is claimed by Google, Georgia Tech or the Wild Dolphin Project.
The model is available to researchers nowGoogle DeepMind’s model page, August 2026: currently in development, will be openly available. No Kaggle page, no Google-published weights on Hugging Face, no entry in Google’s own Gemma variant list.
Dolphins have namesSupported, and not by this model. King and Janik, PNAS 2013: dolphins respond to synthetic copies of their own signature whistle and not to other dolphins’ whistles.
A dolphin used the word for sargassumA dolphin mimicked the synthetic whistle. The Wild Dolphin Project’s own note: nothing to indicate that this word was used in context.
Four claims from the coverage, checked against primary sources.

Where it actually stands, in August 2026

Google said in April 2025 that it planned to share DolphinGemma as an open model “this summer.” It did not. As of this week, the DeepMind model page describes it as “currently in development” and says it “will be openly available” on release. There is no page for it on Kaggle, no Google-published weights on Hugging Face, and no entry in Google’s own list of Gemma variants. Several secondary sites state confidently that the model is downloadable from all three. We checked each one. It is not.

No peer-reviewed results from the model have been published either. The Wild Dolphin Project’s public field-season blogs for 2026 describe encounters, photo-identification work and two pregnant females — not AI findings.

The field itself has not stood still. In June 2026 an independent team published Dolph2Vec, a self-supervised model trained on more than five years of recordings, which they report outperforms general-purpose audio models at classifying and detecting signature whistles. Notably, what they claim is structure and interpretable acoustic units — not meaning, and not translation.

So: an unreleased model, a forty-year field study, a synthetic-whistle experiment that has produced one carefully-caveated mimic, and steady, real progress on finding structure in a signal nobody can yet read. That is a good story. It is simply not the one that says we can talk to dolphins.

The transferable lesson is the one from the demo above. Every time you are told an AI system “understands” something — your customers, your market, a disease, another species — the useful question is whether it has demonstrated understanding or demonstrated prediction. Those are very different claims, they get reported in identical language, and only one of them tells you why.