Ornith 1.0 9B MTP — GGUF (llama.cpp speculative decoding) GGUF builds of deepreinforce ai/Ornith 1.0 9B with the KL distilled MTP draft head from protoLabsAI/Ornith 1.0 9B MTP baked into the trunk — llama.cpp does lossless multi token self speculative decoding out of the box, no separate draft model to wire up. Every file here carries the nextn head, so spec type draft mtp just works. Two things to know before you pick a file: Ampere and older: use Q4 K M . It is smaller and faster than everything else here. Blackwell (RTX 50xx / PRO 6000): use NVFP4 . MTP compounds with NVFP4's tensor core GEMMs where it only partially helps K quants — so NVFP4 is the fastest rung on this hardware by ~28%, despite being a hair larger than Q4 K M . The measured mechanism is below. Want the base with no MTP head? deepreinforce ai/Ornith 1.0 9B GGUF . Files File Size Form Use : Ornith 1.0 9B MTP NVFP4.gguf 6.6 GB bundled Blackwell: fastest rung (306 tok/s +MTP) Ornith 1.0 9B MTP Q8 0.gguf 9.8 GB bundled reference quality / largest relative MTP gain Ornith 1.0 9B MTP Q6 K.gguf 7.6 GB bundled near lossless quant Ornith 1.0 9B MTP Q5 K M.gguf 6.6 GB bundled balanced quality Ornith 1.0 9B MTP Q4 K M.gguf…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy