SmoothConv SmoothConv is a high quality Chinese multi channel conversational speech dataset with expert human annotations , developed by ASLP@NPU and QualiaLabs as part of the SmoothConv–DuplexConv corpus family. Companion dataset: DuplexConv on HuggingFace (2,000 hours, LLM assisted annotation). SmoothConv and DuplexConv are constructed from the same underlying conversational sources. SmoothConv provides high fidelity human annotations for benchmarking and supervised training; DuplexConv offers large scale annotations for Speech LLM pre training and data driven modeling. Dataset Overview SmoothConv contains 100 hours of naturally occurring multi party Chinese conversations recorded in multi channel environments across Tutoring and Social Chat scenarios. Unlike corpora dominated by read speech or scripted interactions, it captures realistic conversational dynamics, including overlapping speech, backchannels, interruptions, pauses, and turn transitions. The dataset is manually annotated by trained experts and provides fine grained conversational labels, making it suitable for turn taking modeling, overlap and interruption detection, full duplex spoken dialogue systems, conversationa…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy