MOSS TTS Family Overview MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity , high‑expressiveness , and complex real‑world scenarios , covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS. Introduction When a single piece of audio needs to sound like a real person , pronounce every word accurately , switch speaking styles across content , remain stable over tens of minutes , and support dialogue, role‑play, and real‑time interaction , a single TTS model is often not enough. The MOSS‑TTS Family breaks the workflow into five production‑ready models that can be used independently or composed into a complete pipeline. MOSS‑TTS : The flagship production model featuring high fidelity and optimal zero shot voice cloning. It supports long speech generation , fine grained control over Pinyin, phonemes, and duration , as well as multilingual/code switched synthesis . MOSS‑TTSD : A spoken dialogue generation model for expressive, multi speaker, and ultra long dialogues. The new v1.0 version a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy