Introduction Welcome to MLAAD: The Multi Language Audio Anti Spoofing Dataset a dataset to train, test and evaluate audio deepfake detection. See the paper for more information. License: Starting from MLAADv8, this dataset will be published under a non commercial license (CC BY NC 4.0). If you want to use this dataset for commercial purposes, you need to either: use a previous version (MLAAD v1 v7) contact us to obtain a commercial license (nicolas.mueller@aisec.fraunhofer.de) Download the dataset Option 1: Hugging Face datasets library Install the datasets package: Login with your Hugging Face account: Then load the dataset in Python: This will automatically handle authentication and download. Option 2: Git + git lfs If you prefer to clone with git, you must first login via Hugging Face: Then clone: Structure The dataset is based on the M AILABS dataset. MLAAD is structured as follows: The file 'meta.csv' contains the following identifiers. For more in these, please see the paper and our website. where reference speaker may or not be present this key has been introduced only in v8. Proposed Usage We suggest to use MLAAD either as new out of domain test data for existing anti spoof…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy