Agents A1 NVFP4/MXFP4 GGUF with MTP Quantized GGUF files for InternScience/Agents A1, a Qwen3.5 35B A3B MoE multimodal model. Attribution Original model : InternScience/Agents A1 — a Qwen3.5 35B A3B MoE multimodal derivative MTP heads : Grafted from unsloth/Qwen3.6 35B A3B MTP GGUF Quantization framework : llama.cpp with NVFP4/MXFP4 MoE routing patches (commits 1f8583c6b , 6e761dbc9 ) Tensor mapping : Follows Unsloth's gold standard MoE quantization strategy Quantized by : s batman Files File Format Size BPW Tensor Type Notes agents a1 NVFP4 MTP.gguf NVFP4 (type 40) 19.8 GB 4.80 80 NVFP4 + 320 Q8 0 + 308 F32 + 5 MTP donor Blackwell native FP4, MTP grafted agents a1 MXFP4 MTP.gguf MXFP4 (type 39) 18.9 GB 4.57 80 MXFP4 + 320 Q8 0 + 308 F32 + 5 MTP donor MXFP4, MTP grafted mmproj agents a1 f16.gguf F16 858 MB 16 334 F16 tensors Vision tower (mmproj), never quantized Tensor Type Mapping Follows Unsloth's gold standard MoE quantization mapping: Tensor Category Count Quant Format 3D expert tensors ( ffn exps ) 80 NVFP4 / MXFP4 Sensitive 2D (attn, embeddings, output, shared experts) 312 Q8 0 Routers ( ffn gate inp , ffn gate inp shexp ) 2 F32 Norms, biases 301 F32 MTP block (blk.40. ) 20…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy