Agents A1 GGUF Quants High quality GGUF quantizations of InternScience/Agents A1, a 35B Qwen3.5 MoE agent model. These files were produced from the BF16 Hugging Face checkpoint with a patched llama.cpp build that supports the qwen35moe architecture. The calibration pass used an importance matrix built from coding/instruction chat data, then each quant was benchmarked against the BF16 GGUF reference. Recommended Files Use case File Notes Best small general purpose quant agents a1 IQ4 XS.gguf Strong quality for size, broad llama.cpp compatibility. Best single user MTP throughput agents a1 IQ4 XS MTP graft headQ6.gguf IQ4 XS body with Q6 K MTP block; measured 1.22x over target only in c1/128 chat serving. Highest MTP acceptance in this run agents a1 Q4 K M MTP graft headQ6.gguf with SPEC DRAFT N MAX=1 91.46% draft acceptance while still 1.15x over target only. Safer MTP quality step up agents a1 Q5 K M MTP graft headQ6.gguf Q5 K M body with the same integrated Q6 K/F32 MTP block; structurally validated, not re profiled yet. Vision / image input for Q4+ quants mmproj agents a1 bf16.gguf Shared BF16 Qwen3VL mmproj for IQ4 XS, Q4 K M, Q5 K M, Q6 K, Q8 0, NVFP4, and the MTP variants. Fast…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy