Agents A1 — GGUF with MTP speculative decoding llama.cpp builds of InternScience/Agents A1 with an MTP draft head grafted in — A1 shipped without the mtp. tensors its Qwen3.5 35B A3B base carries, so no other A1 GGUF can do spec type draft mtp . This one can: +46% measured, no separate draft model. Files file size notes Agents A1 MTP NVFP4.gguf 20.8 GB NVFP4 experts/attention, Q8 0 trunk, MTP head Agents A1 MTP Q8 0.gguf 37.8 GB reference quality, MTP head Measured (RTX A6000 48 GB, n 150 greedy) config gen tok/s NVFP4, no spec 127.7 NVFP4, spec type draft mtp 187.0 (+46%, draft acceptance 0.62) A 35B class agentic MoE at 187 tok/s on a prosumer card in ~21 GB. The MTP head was grafted from the base Qwen3.5 35B A3B — never re trained on A1 — and acceptance holds anyway. Blackwell (RTX PRO 6000 / RTX 50xx class), MTP on, mean of 6 diverse prompts: NVFP4 + draft mtp: 305 tok/s (287–336) A 35B A3B at the same speed our 9B runs — the NVFP4×MTP multiplication (verify step batching feeds the FP4 tensor cores) reproduces on MoE. Finding writeup on the protoLabs Ornith cards. Runtime compatibility llama.cpp (spring 2026+), LM Studio, and recent Ollama (~0.31+): ✅. Older Ollama fails with "…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy