Model Card for Mimi Mimi codec is a state of the art audio neural codec, developped by Kyutai, that combines semantic and acoustic information into audio tokens running at 12.5Hz and a bitrate of 1.1kbps. Model Details Model Description Mimi is a high fidelity audio codec leveraging neural networks. It introduces a streaming encoder decoder architecture with quantized latent space, trained in an end to end fashion. It was trained on speech data, which makes it particularly adapted to train speech language models or text to speech systems. Developed by: Kyutai Model type: Audio codec Audio types: Speech License: CC BY Model Sources Repository: repo Paper: paper Demo: demo Uses How to Get Started with the Model Usage with transformers Use the following code to get started with the Mimi model using a dummy example from the LibriSpeech dataset (~9MB). First, install the required Python packages: Then load an audio sample, and run a forward pass of the model: Usage with Moshi See the main README file. Direct Use Mimi can be used directly as an audio codec for real time compression and decompression of speech signals. It provides high quality audio compression and efficient decoding. Out…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy