Ornith Agents A1 3.6 35B A3B GGUF Multi quant GGUF with a grafted MTP block for native speculative decoding in llama.cpp. This repository contains GGUF quantizations of tepirale/Ornith Agents A1 3.6 35B A3B dare ties , a dare ties merge on top of Qwen3.5 35B A3B (MoE), with a Multi Token Prediction (MTP) block grafted in from the official Qwen3.6 35B A3B GGUF. This lets the model use its own MTP block as a draft model for speculative decoding, with no need for a separate external draft model. 💡 TL;DR: download a quant + the MTP block, launch llama server with spec type draft mtp , and get native speculative decoding with ~45% draft acceptance and up to ~260 tok/s measured on an H100 (25 GB VRAM in use). GPU NVIDIA RTX 6000 Ada Generation 96GB V RAM Table of contents Available quants MTP graft — technical details Usage (llama.cpp) basic code and Performance benchmark Base model and lineage Limitations and notes Available quants Quant Size Draft head (MTP) precision Recommended use Q4 K M 20.5 GiB Q8 0 General use / limited VRAM Q5 K M 23.9 GiB Q8 0 Quality/size balance Q6 K 27.4 GiB Q8 0 High fidelity Q8 0 35.2 GiB Q8 0 Maximum fidelity f16 69.4 GiB F16 Reference / debugging Each q…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy