Multilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC style classification. Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time shifting, noise injection, pitch variation). The dataset is organized in folders per command label. Metadata files are included to facilitate training and evaluation. Languages English Kazakh Russian Tatar Arabic Turkish French German Spanish Italian Catalan Persian Polish Dutch Kinyarwanda Structure The dataset includes: One folder per speech command (e.g., yes/ , no/ , go/ , stop/ , etc.) Metadata files: training list.txt validation list.txt testing list.txt label map.json lang map.json License All audio samples are redistributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, as permitted by the original sources. Please cite the original datasets below if you use this dataset in yo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy