natori irodori tts dataset High quality single speaker Japanese speech dataset prepared for Irodori TTS LoRA training from multiple long form さなちゃんねる videos featuring 名取さな. Dataset Summary This repository is a merged export of four curated subsets derived from public YouTube playlists on さなちゃんねる . Current top level merged export: 40,995 total utterances 33,398 training examples 7,597 validation examples 43.4277 hours of speech speaker id: natori sana The merged dataset currently contains these subsets: paranormasight 2026 : 5.6632h chats 2024 : 16.4336h chats 2025 : 15.5682h chats 2026 : 5.7628h Source Data All source media comes from the YouTube channel さなちゃんねる . Playlists used in the current release: Paranormasight gameplay/commentary: https://youtube.com/playlist?list=PLPDt0GwV6dYdk79D0obyy0IOtFogXB 6R&si=0n2JxJatr8dtdaiU 2026 chat streams: https://youtube.com/playlist?list=PLPDt0GwV6dYdx80xch7a9IKpM6leRn8XJ&si=1fOo79y2 aImnq8R 2025 chat streams: https://youtube.com/playlist?list=PLPDt0GwV6dYdvClZPJOS9aupubucXm6nh&si=wJNgnCbt4Qp79v48 2024 chat streams: https://youtube.com/playlist?list=PLPDt0GwV6dYfUJrt3deX9yoisWDlJ2pUk&si=nQf7U aRhMrMgvJ8 Source format: long form gameplay comme…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy