Ornith 1.0 35B MTP (GGUF) ๐ก Strix Halo / gfx1151? A fork only ROCmFP4 build (ROCmFPX, ~19 GiB, Vulkan 86.7 t/s MTP) is available separately: โ singulared/Ornith 1.0 35B MTP ROCmFP4 GGUF Ornith 1.0 35B (DeepReinforce) with an embedded MTP (Multi Token Prediction / nextn ) head grafted in , enabling self speculative decoding in llama.cpp at identical output quality . Which file? file size decode t/s acceptance notes ornith 1.0 35b MTP Q4 K M.gguf 20.6 GiB ~80 0.847 recommended ornith 1.0 35b MTP Q8 0.gguf 35.2 GiB ~63โ66 0.859 near lossless weights Sizes are weights only โ budget additional headroom for the KV cache and compute buffers, which grow with context length. At long context (128K) plan well above the file size. Both carry the MTP head at Q8 0 precision . That matters: the nextn.eh proj tensor is only ~9 MB but it largely determines draft acceptance โ quantizing it down costs a substantial slice of the speedup, so it is kept at Q8 even in the Q4 K M build. Why Ornith 1.0 35B is an agentic coder fine tuned from Qwen3.6 35B A3B (same qwen35moe architecture, same tokenizer). The base Qwen ships with an embedded MTP head; Ornith's release does not โ so it decodes without self sโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy