Qwen3.5 122B A10B NVFP4 MTP GGUF NVFP4 GGUF of Qwen/Qwen3.5 122B A10B with the MTP (Multi Token Prediction) head retained. Structurally identical to the sibling artifact Incarnas/Qwen3.5 122B A10B NVFP4 GGUF plus one additional layer block carrying the MTP draft weights, enabling self speculative decoding via llama.cpp's spec type draft mtp path. 122B total parameters, ~10B active per token via 256 MoE experts (8 routed plus 1 shared). 48 base transformer blocks plus 1 MTP block (block index 48). 77 GiB main GGUF, 871 MiB mmproj sidecar carrying the vision tower. For reproducibility / audit / debugging details — full bug discovery narrative, per test bench tables, hardware specs, failure modes encountered during this run, and step by step reproduction recipe — see the companion TECHNICAL REPORT.md (or the PDF version on the methodology repo). The full methodology toolchain — the converter patch, bench scripts, raw bench JSONs, power telemetry CSVs, and gsm8k samples — lives at bit incarnas/nvfp4 mtp conversions (v1.0). Updates 2026 05 18 — The converter patch required by this release has been merged upstream as ggml org/llama.cpp 23237 (master commit 1867a0c69 , first contained in…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy