Model Card for Moshi Moshi is a speech text foundation model and full duplex spoken dialogue framework Model Details Candle version (Rust) with 8 bits quantization. Model Description Moshi is a speech text foundation model that casts spoken dialogue as speech to speech generation. Starting from a text language model backbone, Moshi generates speech as tokens from the residual quantizer of a neural audio codec, while modeling separately its own speech and that of the user into parallel streams. This allows for the removal of explicit speaker turns, and the modeling of arbitrary conversational dynamics. Moshi also predicts time aligned text tokens as a prefix to audio tokens. This “Inner Monologue” method significantly improves the linguistic quality of generated speech and provides streaming speech recognition and text to speech. As a result, Moshi is the first real time full duplex spoken large language model, with a theoretical latency of 160ms, 200ms in practice. Developed by: Kyutai Model type: Multimodal speech text foundation model Language(s) (NLP): English License: CC BY Model Sources Repository: repo Paper: paper Demo: demo Uses Direct Use The model can be used as a convers…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy