A new version of streaming Sortformer v2.1 has been released, providing greater robustness for meeting speech. Streaming Sortformer Diarizer 4spk v2 img { display: inline; } This model is a streaming version of Sortformer diarizer. Sortformer[1] is a novel end to end neural model for speaker diarization, trained with unconventional objectives compared to existing end to end diarization models. Streaming Sortformer[2] employs an Arrival Order Speaker Cache (AOSC) to store frame level acoustic embeddings of previously observed speakers. Sortformer resolves permutation problem in diarization following the arrival time order of the speech segments from each speaker. This speaker diarization model can be used to enable the NeMo Voice Agent to recognize speakers in conversations. See the NeMo Voice Agent and the YAML configuration for more details. Discover more from NVIDIA: For documentation, deployment guides, enterprise ready APIs, and the latest open models—including Nemotron and other cutting edge speech, translation, and generative AI—visit the NVIDIA Developer Portal at developer.nvidia.com. Join the community to access tools, support, and resources to accelerate your development…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy