🗣️ Lipi Ghor লিপিঘর — Bengali Speech Dataset (bn 882 SSTT) green) Lipi Ghor (লিপিঘর, meaning "House of Scripts" ) is a large scale Bengali speech dataset designed for automatic speech recognition (ASR), speaker diarization, and spoken language research. It is one of the largest open Bengali speech corpora with aligned speaker, transcription, and timestamp annotations. Built by Team Villagers as part of DL Sprint 4.0 . Dataset Details Dataset Description Lipi Ghor is a large scale, multi domain Bengali speech corpus covering ~882 hours of audio sourced from 1,019 YouTube videos across 596 unique channels. Each video has been processed through speaker diarization (pyannote audio) and aligned with Bengali caption transcripts to produce structured SSTT (Speaker, Speech, Transcription, Timestamp) annotations. The dataset covers a wide range of spoken Bengali domains, registers, and regional dialects, making it one of the most diverse open Bengali speech resources available. Curated by: Team Villagers — Sanjid Hasan, A H M Fuad, Risalat Labib, Bayazid Hasan Competition: DL Sprint 4.0 Language(s): Bengali / Bangla ( bn ) License: CC BY 4.0 Dataset Sources Repository: Sanjidh090/Lipi Ghor…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy