MioCodec 25Hz 24kHz: Lightweight Neural Audio Codec for Efficient Spoken Language Modeling MioCodec 25Hz 24kHz is a lightweight and fast neural audio codec designed for efficient spoken language modeling. Based on the Kanade Tokenizer implementation, this model features an integrated wave decoder (iSTFTHead) that directly synthesizes waveforms without requiring an external vocoder. For higher audio fidelity at 44.1 kHz, see MioCodec 25Hz 44.1kHz. 🌟 Overview MioCodec decomposes speech into two distinct components: 1. Content Tokens: Discrete representations that primarily capture linguistic information and phonetic content ("what" is being said) at a low frame rate (25 Hz). 2. Global Embeddings: A continuous vector representing broad acoustic characteristics ("how")—including speaker identity, recording environment, and microphone traits. By disentangling these elements, MioCodec is ideal for Spoken Language Modeling . Key features Lightweight & Fast: Integrated wave decoder (iSTFTHead) enables direct waveform synthesis without an external vocoder. Ultra Low Bitrate: Achieves high fidelity reconstruction at only 341 bps . End to End Design: Single model architecture from audio inpu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy