Russian Podcasts (unlabeled) ~186k unlabeled Russian language podcast episodes scraped from the web, packaged as Parquet shards with the audio bytes embedded. The audio has no transcripts — this is an unsupervised / self supervised audio corpus, suitable for ASR pre training, speech representation learning, TTS data mining, audio classification, and similar tasks. Each row contains: audio — the podcast episode (MP3, mostly 128 kbps / 44.1 kHz stereo), decoded on the fly via the Audio feature. a set of flat scalar metadata columns: id , name , lang , duration sec , audio bytes , rating , rating votes , rating scores , read count , reviews count , citations count , author / narrator names, publisher name , genres names , tags csv , series name , written dt , updated at , annotation plain , annotation html , and a few more. raw metadata json — the full original metadata record serialized as a string. Loading examples The audio was scraped from publicly reachable sources on the internet and is provided as is for research use.
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy