MOSS TTS Family MOSS TTS Local Transformer v1.5 MOSS TTS Local Transformer v1.5 is continued from MOSS TTS Local Transformer v1.0. It preserves the main 1.0 capabilities, including zero shot voice cloning, long form speech generation, token level duration control, Pinyin/IPA pronunciation control, multilingual synthesis, and code switching. For the full 1.0 feature walkthrough, input schema, and evaluation tables, please refer to the MOSS TTS Local Transformer v1.0 README. Compared with MOSS TTS Local Transformer v1.0, v1.5 focuses on the following improvements: Higher fidelity stereo audio modeling : v1.5 uses MOSS Audio Tokenizer v2 as the audio tokenizer, supporting native 48 kHz stereo input and output for richer spatial detail and more natural perceived audio quality. Since the codec output is stereo, save the [channels, samples] tensor returned by processor.decode(...) directly. Stronger multilingual synthesis with language tags : when the language field is omitted, v1.5 may improve some languages and regress slightly on others compared with 1.0. When the language is specified, v1.5 is stronger than 1.0 on almost all supported languages. Set the tag w…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy