WILDELDER: A CHINESE ELDERLY SPEECH DATASET FROM THE WILD WITH FINE GRAINED MANUAL ANNOTATIONS Paper: https://huggingface.co/papers/2510.09344 Code: https://github.com/NKU HLT/WildElder WildElder is a speech dataset focused on elderly scenarios. It contains raw audio and corresponding text annotations and can be used for ASR, speaker related tasks, and front /back end speech processing research. The data was collected and cleaned from real world environments to preserve diversity and realistic noise conditions. Viewer Configuration This repository is configured for the Hugging Face Dataset Viewer using split level CSV metadata files: train.csv validation.csv test.csv file name : relative path to the audio file in the repository text : transcription utt id : utterance ID from the original Kaldi style manifest source path : original path recorded in wav.scp The Dataset Viewer uses these CSV files and resolves the linked files from the existing audio/ directory, so the corpus can be browsed without duplicating audio under new split folders. Loading Notes validation is generated from the original data split/dev manifest. The Dataset Viewer is configured to use the extracted audio/ tree…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy