(简体中文 English 日本語) Introduction github repo : https://github.com/FunAudioLLM/SenseVoice SenseVoice is a speech foundation model with multiple speech understanding capabilities, including automatic speech recognition (ASR), spoken language identification (LID), speech emotion recognition (SER), and audio event detection (AED). [//]: ( ) Homepage | What's News | Benchmarks | Install | Usage | Community Model Zoo: modelscope, huggingface Online Demo: modelscope demo, huggingface space Highlights 🎯 SenseVoice focuses on high accuracy multilingual speech recognition, speech emotion recognition, and audio event detection. Multilingual Speech Recognition: Trained with over 400,000 hours of data, supporting more than 50 languages, the recognition performance surpasses that of the Whisper model. Rich transcribe: Possess excellent emotion recognition capabilities, achieving and surpassing the effectiveness of the current best emotion recognition models on test data. Offer sound event detection capabilities, supporting the detection of various common human computer interaction events such as bgm, applause, laughter, crying, coughing, and sneezing. Efficient Inference: The SenseVoice Small mo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy