Bagpiper Pretraining Data Bagpiper Pretraining Data is the public rich captioned audio snapshot associated with Bagpiper, an open ended audio language model that learns bidirectional mappings between audio and comprehensive text descriptions across speech, music, environmental sound, and mixtures. The en metadata describes the primary rich caption language. Source audio can contain speech or singing in other languages; it is not an English only audio guarantee. The repository contains 155,151,789 rows in 7,780 valid Parquet shards across 18 source family directories, using 5.229 TB. It is intended for large scale research workflows; select only the source families needed for your experiment rather than downloading the entire repository by default. Repository snapshot Source family directory Rows Valid Parquet shards GB : : : audiocaps 35,206 2 1.217 audioset 1,346,522 68 47.973 clotho aqa 4,193 1 0.130 clotho train 29,434 15 0.905 emilia en 15,664,702 784 528.702 fma 2,317,679 116 84.010 laion audio 300m part1 22,439,013 1,122 469.972 laion audio 300m part2 23,840,302 1,193 527.090 laion audio 300m part3 24,424,316 1,222 541.761 laion audio 300m part4 18,983,198 950 352.998 laion c…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy