community 1 speaker diarization This pipeline ingests mono audio sampled at 16kHz and outputs speaker diarization. stereo or multi channel audio files are automatically downmixed to mono by averaging the channels. audio files sampled at a different rate are resampled to 16kHz automatically upon loading. The main improvements brought by Community 1 are: improved speaker assignment and counting simpler reconciliation with transcription timestamps with exclusive speaker diarization easy offline use (i.e. without internet connection) (optionally) hosted on pyannoteAI cloud Setup 1. pip install pyannote.audio 2. Accept user conditions 3. Create access token at hf.co/settings/tokens . Quick start Benchmark Out of the box, Community 1 is much better than speaker diarization 3.1 . We report diarization error rates (in %) on large collection of academic benchmarks (fully automatic processing, no forgiveness collar, nor skipping overlapping speech). Benchmark (last updated in 2025 09) legacy (3.1) community 1 precision 2 AISHELL 4 12.2 11.7 11.4 AliMeeting (channel 1) 24.5 20.3 15.2 AMI (IHM) 18.8 17.0 12.9 AMI (SDM) 22.7 19.9 15.6 AVA AVD 49.7 44.6 37.1 CALLHOME (part 2) 28.5 26.7 16.6 DIHA…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy