Fun ASR MLT Nano 2512: transcribe.cpp GGUF GGUF conversions of FunAudioLLM/Fun ASR MLT Nano 2512 for use with transcribe.cpp. Ported from upstream commit cf67a938bf2829959d08fdfb84e186eff02a67ff, pinned 2026 05 06. Validated against the FunASR reference at transcribe.cpp commit f094d28 on 2026 05 06. Offline speech to text covering 31 languages, with focused optimization on East and Southeast Asian languages: Chinese, English, Cantonese, Japanese, Korean, Vietnamese, Indonesian, Thai, Malay, Filipino, plus Arabic, Hindi, and 19 European languages (Bulgarian, Croatian, Czech, Danish, Dutch, Estonian, Finnish, Greek, Hungarian, Irish, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian, Swedish). Same architecture as Fun ASR Nano 2512 (~800M trainable parameters: frozen SenseVoiceEncoderSmall + 2 layer audio adaptor + bundled Qwen3 0.6B LLM); trained on a smaller multilingual corpus ("hundreds of thousands of hours" per the model card, vs Nano's "tens of millions"). Takes a 16 kHz mono WAV and emits text. Not streaming, no translation, no timestamps. ITN (inverse text normalization) is supported by the model and exposed via the itn CLI flag and transcribe fu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy