MIA LALM Audio datasets for Membership Inference Attacks against Large Audio Language Models (arXiv, code). The audio is distributed as one zstd compressed tar per dataset family. A few large archives download far faster than ~90k individual WAV files (no per file overhead, no HTTP 429), and each extracts to the exact layout the attack runners expect. Download With the code repository cloned, download each family into MIA on dataset/data/ and extract it there — the loader picks it up automatically, no environment variable needed: Each dataset then lives under MIA on dataset/data/ mia dataset/ , matching the manifest paths the runners expect (e.g. voxpopuli mia dataset/tts 61 2/tts dataset.json ). To keep the data elsewhere, extract anywhere and export AUDIO MIA DATA ROOT=/that/path . Layout Each archive expands to mia dataset/... containing audio plus the JSON/CSV manifests. The same manifests are also available loose in this repo for browsing.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy