Llama Mimi: Speech Language Models with Interleaved Semantic and Acoustic Tokens 🤗 HuggingFace 📄 Paper 🗣️ Online Demo 🧑💻 Code Introduction Llama Mimi is a speech language model that uses a unified tokenizer (Mimi) and a single Transformer decoder (Llama) to jointly model sequences of interleaved semantic and acoustic tokens. Trained on ~240k hours of English audio, Llama Mimi achieves state of the art performance in acoustic consistency on SALMon and effectively preserves speaker identity. Visit our demo site to hear generated speech samples. Models Models 🤗 Hugging Face Llama Mimi 1.3B llm jp/Llama Mimi 1.3B Llama Mimi 8B llm jp/Llama Mimi 8B How to Use Install dependencies: Generate audio continuations from a given audio prompt: Pretraining & Evaluation of Llama Mimi Check out our repository: https://github.com/llm jp/llama mimi Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy