S2 Pro — GGUF ALPHA — EXPERIMENTAL The inference engine (s2.cpp) is an early stage, community built project. Expect rough edges and breaking changes. Not production ready. GGUF quantized weights of Fish Audio S2 Pro, a high quality multilingual text to speech model with voice cloning support, packaged for local inference with s2.cpp — a pure C++/GGML engine with no Python dependency. License: Fish Audio Research License — free for research and non commercial use. Commercial use requires a separate license from Fish Audio. See LICENSE.md and fish.audio. Files File Size s2 pro f16.gguf 9.3 GB s2 pro q8 0.gguf 5.3 GB s2 pro q6 k.gguf 4.3 GB s2 pro q5 k m.gguf 3.8 GB s2 pro q4 k m.gguf 3.4 GB s2 pro q3 k.gguf 2.9 GB s2 pro q2 k.gguf 2.4 GB tokenizer.json 12 MB All GGUF files contain both the transformer weights and the audio codec in a single file. Requirements GPU with Vulkan support (AMD/NVIDIA/Intel) or CPU with enough RAM s2.cpp built from source (C++17 + CMake) VRAM guide VRAM Recommended ≥ 8 GB q8 0 6–8 GB q6 k 4–6 GB q5 k m 3–4 GB q4 k m < 3 GB q3 k / q2 k (quality degrades) CPU only q4 k m or lower (slow) Quick start Voice cloning Reference audio: 5–30 seconds, clean recording,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy