YouTube ASR Caption Dataset (Cantonese) This dataset was built from YouTube videos with manually provided captions in Cantonese. We used SenseVoice to re transcribe the audio and filtered segments to build a high quality collection of audio caption pairs. What’s included Segments where the ASR output is identical to the original caption — likely clean. Segments where differences are only homophones (同音字) or English words — likely ASR mistakes. This combination… See the full description on the dataset page: https://huggingface.co/datasets/ming030890/youtube caption yue.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy