MioCodec 25Hz 44.1kHz v2: Lightweight Neural Audio Codec for Efficient Spoken Language Modeling MioCodec 25Hz 44.1kHz v2 is an upsampled, high fidelity version of the MioCodec 25Hz 24kHz model. By integrating an UpsamplerBlock inspired by Inworld TTS 1 into the decoder, this model reconstructs 44.1 kHz audio from the standard 25 Hz token stream. 🌟 What's New in v2 This model is a fine tuned version of MioCodec 25Hz 24kHz with the following architectural enhancements: 44.1 kHz Output: Achieves higher audio fidelity compared to the base 24 kHz model. UpsamplerBlock + SnakeBeta: We adopted the UpsamplerBlock architecture from Inworld TTS 1 and enhanced it by integrating SnakeBeta activations. This combination allows the decoder to effectively predict and generate high frequency components, enabling clear 44.1 kHz reconstruction from the lower resolution input. Token Compatibility: During fine tuning, the content branch was frozen. This means the discrete tokens generated by this model are identical to those from MioCodec 25Hz 24kHz . You can take any TTS model trained on the 24kHz tokens and simply swap the codec to this v2 model during inference to instantly upgrade the audio qualit…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy