MOSS TTS Family Overview MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity , high‑expressiveness , and complex real‑world scenarios , covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS. The model architecture and tokenizer are detailed in the paper MOSS Audio Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models. Introduction When a single piece of audio needs to sound like a real person , pronounce every word accurately , switch speaking styles across content , remain stable over tens of minutes , and support dialogue, role‑play, and real‑time interaction , a single TTS model is often not enough. The MOSS‑TTS Family breaks the workflow into five production‑ready models that can be used independently or composed into a complete pipeline. MOSS‑TTS : MOSS TTS is the flagship production TTS foundation model, centered on high fidelity zero shot voice cloning with controllable long form synthesis, pronunciation, and multilingual/code switched speech. It serves as the…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy