Gepard GE nerative, P rosody aware, A utoregressive text to speech model for R ealtime D ialogue Gepard is a text to speech model built for real time conversation. It starts speaking the moment text begins arriving, generating audio piece by piece instead of waiting for a full sentence — so it feels like a live voice, not a recording. It's a single language model that learned text and speech together, so the output carries natural rhythm and timing rather than the flat, stitched tone of older pipelines. The name evokes "Gepard"( /geh PART/ ), German for cheetah — a nod to the model's low latency, high throughput streaming. Want to use Gepard in production without hosting it yourself? You can skip the deployment and use our real time TTS API — a fully managed, Cartesia compatible service with millisecond time to first chunk, voice cloning, and streaming built in. If you've used Cartesia, you already know how to use it: point the base URL at https://api.nineninesix.ai and the official Cartesia SDKs just work. Roughly 22 hours of audio for $5, and free credits to start (no credit card required). Try the live demo or grab an API key. Highlights: One clean pass per frame — the whole aud…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy