A reproducible, fully non-invasive pipeline for killer-whale (Orcinus orca) call structure, behavioural context, and playback-response analysis from public recordings.
One frozen audio foundation model (AVES2) runs unchanged across the primary analysis; the contribution is the rigour around it — measuring the recording-site shortcut that fools naive probes, reporting only what survives explicit null baselines, and checking the two primary representation results under a second frozen encoder.
A frozen model “decodes ecotype” at 0.91 pooled — but hold out a whole recording site and it collapses to 0.23 (chance 0.25), while the site itself decodes at 0.948. The confound is measured instead of hidden.
Within five fixed recording sites the embeddings still separate ecotypes; the hydrophone is held constant, so this is not a site shortcut.
Catalogue call types are recovered within a site (14 SRKW types 0.71; 18 NRKW types 0.97) and — unlike ecotype — transfer to an independent site (0.64; 0.83 on unambiguous types): the control ecotype failed.
On animal-borne tags, communicative calls decode movement-defined context (foraging vs not 0.77; three-way 0.58) with every individual held out, and specific call types occur in specific contexts (V=0.40). Controls rule out call rate and loudness as the explanation.
Calls don’t come at random: one call predicts the next, and Southern-Resident sequences carry information a first-order model can’t explain (candidate phrase S01→S04→S01). A prerequisite for combinatorial coding — not meaning.
Re-analysis of a published playback: free-ranging whales replied to same-pod broadcasts (6/6) and stayed silent to different-pod ones (0/6), naive animals — and the model recovers the dialect call-type space underlying the response contrast (purity 0.439 vs 0.05 null). The experiment is prior work; the response tracks dialect, not content.
Frozen NatureLM-audio repeats the site-controlled SRKW and NRKW call-type checks above chance, preserves VFPA→SMRU transfer (0.682 vs 0.20 chance), and separates FEROP K-type exemplars (purity 0.366 vs 0.050 chance). Cross-encoder check, not semantic meaning.
Generated by the scripts in the repository; full metrics live in reports/ and reproduction status in the environment manifest.










Use the model-facing console to select or upload a candidate signal, inspect call-type/dialect structure, production-context evidence, playback-response design, and a field-planning readout from the same claim-bounded stack.
Configure a controlled study, produce a randomized encounter allocation, and export a protocol, catalogue audit, manifest, operator sheet, and genuinely blinded scoring sheet. The website does not create field-release audio.
Review-kit JSON includes the protocol, condition table, encounter allocation, manifest, blind scoring schema, and release checklist. It contains no field-release audio.
Archival recordings can characterise call structure and its confounds. What an archive cannot do — even in principle — is measure how a whale responds to a signal deliberately broadcast to it. That needs a controlled in-situ playback with a permitted field collaborator. The analysis and review tooling are built; target-specific stimuli, calibration, authorization, and field execution are not.
Clone it, install the analysis extras, run the tests, and reproduce the playback-response statistic from the committed trial table:
git clone https://github.com/ladyFaye1998/OrcaDolittle
cd OrcaDolittle
python -m pip install -e ".[dev,analysis]"
python -m pytest -q
python scripts/run_playback_response_stats.py
python scripts/build_field_playback_package.py --config field_playback/field_config.example.json --output field_playback/orca_playback_review_package.zip
The example build is review-only and contains no broadcast audio. See the controlled-playback protocol and release boundary.
For permitted playback or biologging studies, include the study population, permit status, season timing, and response measures.
The project is intentionally acoustic and claim-bounded: it models recorded underwater audio and measured behavioural-response evidence, not the whale’s full perceptual world, not translation, and not semantic meaning. The full limitation register is in the repository.
Orcas may use multimodal or unrecorded cues. The analysis therefore makes acoustic claims only, and the field design pairs hydrophones with visual, group-ID, movement, and response measures.
Killer whales hear across high-frequency ranges, while public recordings come from mixed hydrophones and conversion paths. Provider is exposed as a confound; positive claims are limited to within-site tests or explicit transfer tests.
The published playback shows a receiver response to familiar dialect calls. It does not prove that whales respond to call content; that remains the permitted field playback that isolates content while holding dialect/social familiarity constant.
NatureLM-audio repeats the two primary representation checks, which addresses part of the single-encoder concern. It still does not turn acoustic separability into semantic interpretation.
See limitations and mitigations and data availability for the source-by-source boundary.
A quick look at what’s inside the corpus, right here on the page. Most inputs are public through the linked repositories; raw third-party archives, large native tag files, and model caches remain external, with a frozen derived-artifact package on Zenodo and Git provenance documented in the repo.
Data are public: DCLDE 2026 (NOAA/NCEI), the DTAG archive (Zenodo, CC-BY), and the FEROP catalogue. Code and the full bibliography are in the repository.
Palmer, K. J., et al. (2025). A public dataset of annotated Orcinus orca acoustic signals for detection and ecotype classification. Scientific Data 12:1137.
Szymanski, M. D., et al. (1999). Killer whale hearing: auditory brainstem response and behavioral audiograms. J. Acoust. Soc. Am. 106:1134–1141.
Johnson, M. P., Tyack, P. L. (2003). A digital acoustic recording tag for measuring the response of wild marine mammals to sound. IEEE J. Oceanic Engineering 28:3–12.
Hagiwara, M. (2023). AVES: animal vocalization encoder based on self-supervision. ICASSP 2023.
Chen, S., et al. (2022). BEATs: audio pre-training with acoustic tokenizers. arXiv.
Ford, J. K. B. (1989). Acoustic behaviour of resident killer whales off Vancouver Island, British Columbia. Canadian Journal of Zoology 67:727–745.
Foote, A. D., Osborne, R. W., Hoelzel, A. R. (2008). Temporal and contextual patterns of killer whale call type production. Ethology 114:599–606.
Filatova, O. A., et al. (2015). Cultural evolution of killer whale calls. Behaviour 152:2001–2038.
Yurk, H., et al. (2002). Cultural transmission within maternal lineages: vocal clans in resident killer whales in southern Alaska. Animal Behaviour 63:1103–1119.
Riesch, R., Ford, J. K. B., Thomsen, F. (2008). Whistle sequences in wild killer whales. J. Acoust. Soc. Am. 124:1822–1829.
Filatova, O. A., Fedutin, I. D., Burdin, A. M., Hoyt, E. (2011). Responses of Kamchatkan fish-eating killer whales to playbacks of conspecific calls. Marine Mammal Science 27:E26–E42.
Far East Russia Orca Project (FEROP). Catalogue of Kamchatkan killer whale calls.
Holt, M. M., Tennessen, J. B., et al. (2024). Calibrated movement data and variables. Zenodo, doi:10.5281/zenodo.13308835.
Tennessen, J. B., et al. (2019). Hidden Markov models reveal temporal patterns and sex differences in killer whale behavior. Scientific Reports 9:14951.
Wilson, R. P., et al. (2006). Moving towards acceleration for estimates of activity-specific metabolic rate (ODBA). J. Animal Ecology 75:1081–1090.
Berthet, M., Surbeck, M., Townsend, S. W. (2025). Extensive compositionality in the vocal system of bonobos. Science 388:104–108.
Girard-Buttoz, C., …, Crockford, C. (2025). Versatile use of chimpanzee call combinations promotes meaning expansion. Science Advances 11:eadq2879.
Sharma, P., et al. (2024). Contextual and combinatorial structure in sperm whale vocalisations. Nature Communications 15:3617.
Stowell, D. (2022). Computational bioacoustics with deep learning. PeerJ 10:e13152.
Ghani, B., et al. (2023). Global birdsong embeddings enable superior transfer learning. Scientific Reports 13:22876.
Kershenbaum, A. (2024). Why Animals Talk. Penguin Press.
Sayigh, L. S., et al. (2025). Widespread sharing of stereotyped non-signature whistle types by wild dolphins. bioRxiv.
Selbmann, A., et al. (2026). Aversive behavioural responses of killer whales to sounds of long-finned pilot whales. Scientific Reports 16:4716.
Bowers, M. T., et al. (2018). Selective reactions to different killer whale call categories in two delphinid species. J. Experimental Biology 221:jeb162479.
Miller, P. J. O., et al. (2004). Call-type matching in vocal exchanges of free-ranging resident killer whales. Animal Behaviour 67:1099–1107.
Robinson, D., et al. (2024). NatureLM-audio: an audio-language foundation model for bioacoustics. arXiv:2411.07186.