Raon Speech 9B Demo Technical Report Raon Speech is a 9B parameter speech language model that supports state of the art speech understanding, answering and generation in English and Korean. This model successfully transforms a pre trained LLM into a SpeechLM to both understand and generate speech without compromising its original language capabilities. It trains on millions of hours of English Korean speech text datasets with the following training stages: (1) speech encoder decoder alignment, (2) end to end SpeechLM pre training, and (3) multi reward DPO based post training. Key Features End to End Speech Language Model : 9B parameter multimodal model built on Qwen3 (36 layers, 4096 hidden dim), Qwen3OmniMoeAudioEncoder (24 layers), Mimi codec (32 quantizers), and ECAPA TDNN speaker encoder. Bilingual Support : State of the art speech understanding, answering, and generation in both English and Korean. Multi Task Capabilities : Supports STT (audio → text), TTS (text → audio), TextQA (text + audio → text), and SpeechChat (audio → text) in a single unified model. Speaker Voice Conditioning : TTS with optional speaker reference audio for voice cloning via ECAPA TDNN embeddings. TTS C…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy