VoxCPM2 VoxCPM2 is a tokenizer free, diffusion autoregressive Text to Speech model — 2B parameters , 30 languages , 48kHz audio output, trained on over 2 million hours of multilingual speech data. Highlights 🌍 30 Language Multilingual — No language tag needed; input text in any supported language directly 🎨 Voice Design — Generate a novel voice from a natural language description alone (gender, age, tone, emotion, pace…); no reference audio required 🎛️ Controllable Cloning — Clone any voice from a short clip, with optional style guidance to steer emotion, pace, and expression while preserving timbre 🎙️ Ultimate Cloning — Provide reference audio + its transcript for audio continuation cloning; every vocal nuance faithfully reproduced 🔊 48kHz Studio Quality Output — Accepts 16kHz reference; outputs 48kHz via AudioVAE V2's built in super resolution, no external upsampler needed 🧠 Context Aware Synthesis — Automatically infers appropriate prosody and expressiveness from text content ⚡ Real Time Streaming — RTF as low as ~0.3 on NVIDIA RTX 4090, and ~0.13 accelerated by Nano VLLM 📜 Fully Open Source & Commercial Ready — Apache 2.0 license, free for commercial use Supported Langua…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy