Dataset card for NatureLM audio training Overview NatureLM audio training is a large and diverse audio language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording. For example, for an in the wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio language model trained on this dataset may respond with "Common yellowthroat". It consists of over 26 million audio text pairs derived from diverse sources including animal vocalizations, insects, human speech, music and environmental sounds: Task Dataset Hours Samples CAP WavCaps (Mei et al., 2023) 7,568 402k CAP AudioCaps (Kim et al., 2019) 145 52k CLS NSynth (Engel et al., 2017) 442 300k CLS LibriSpeechTTS (Zen et al., 2019), VCTK (Yamagishi et al. 2019) 689 337k CAP Clotho (Drossos et al. 2020) 25 4k CLS, DET, CAP Xeno canto (Vellinga & Planque, 2015) 10,416 607k CLS, DET, CAP iNaturalist (iNaturalist) 1,539 320k CLS, DET, CAP Watkins (Sayigh et al., 2016) 27 15k CLS, DET Animal Sound Archive (Museum für Naturkunde Berlin) 78 16k DET Sa…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy