Russian Podcasts (unlabeled) ~186k unlabeled Russian language podcast episodes scraped from the web, packaged as Parquet shards with the audio bytes embedded. The audio has no transcripts — this is an unsupervised / self supervised audio corpus, suitable for ASR pre training, speech representation learning, TTS data mining, audio classification, and similar tasks. Each row contains: audio — the podcast episode (MP3, mostly 128 kbps / 44.1 kHz stereo), decoded on the fly via the Audio feature. a set of flat scalar metadata columns: id , name , lang , duration sec , audio bytes , rating , rating votes , rating scores , read count , reviews count , citations count , author / narrator names, publisher name , genres names , tags csv , series name , written dt , updated at , annotation plain , annotation html , and a few more. raw metadata json — the full original metadata record serialized as a string. Loading examples The audio was scraped from publicly reachable sources on the internet and is provided as is for research use.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy