Ornith 1.0 35B — GGUF (llama.cpp, single GPU tp=1 ) Single GPU llama.cpp GGUF package for deepreinforce ai/Ornith 1.0 35B . The supported serving policy is tp=1 — one model copy per GPU. Multi GPU tensor parallel serving is intentionally out of scope for this version. This release ships six body quants (Q3 K M → Q8 0) plus an integrated IQ4 XS MTP graft that adds a native multi token prediction (MTP) draft head for low concurrency speculative decode. Quant quality is measured against the upstream BF16 GGUF with a native llama.cpp next token top 64 KL divergence probe over 32 coding prompts. TL;DR — pick an artifact Use case Recommended artifact Key numbers Default serving speed ornith 1.0 35b Q4 K M.gguf 19.71 GiB on disk, 21.31 GiB loaded VRAM, 243.3 tok/s c1, 655.6 tok/s c16 Lowest memory ornith 1.0 35b Q3 K M.gguf 15.61 GiB on disk, 17.27 GiB loaded VRAM, 240.5 tok/s c1, 493.0 tok/s c16 Middle footprint ornith 1.0 35b IQ4 XS.gguf 17.64 GiB on disk, 19.34 GiB loaded VRAM, 0.1426 mean top 64 KLD nats Highest fidelity / footprint ornith 1.0 35b Q6 K.gguf 26.56 GiB on disk, 28.03 GiB loaded VRAM, 0.0165 mean top 64 KLD nats, 32/32 top 1 Native low concurrency MTP ornith 1.0 35b IQ4…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy