Qwen3.5 122B A10B heretic MTP NVFP4 NVFP4 (W4A4) quantization of trohrbaugh/Qwen3.5 122B A10B heretic. Base Model: Qwen3.5 122B A10B heretic (MoE: 122B total, ~10B active, abliterated with KL ~0.09) Quantization: NVFP4 (W4A4) — weights and activations Size: 76GB (16 shards + MTP shard + visual shard) Quantized: Language backbone MoE expert and attention layers NOT Quantized: Vision encoder, merger, LM head, embed tokens, linear attention, MoE gates, MTP heads (remain BF16) MTP: Working speculative decoding — 785 tensors spliced from base Qwen/Qwen3.5 122B A10B in BF16 Tokenizer: From base Qwen3.5 Usage with vLLM Tested on vLLM 0.19+. MTP Throughput (2x RTX 6000 Pro Blackwell, TP=2) MTP Speculative Tokens tok/s 0 (disabled) ~105 1 ~115 2 ~145 3 ~170 6 ~190 Quantization Quantized with llm compressor (compressed tensors v0.14.1.dev28). Format: nvfp4 pack quantized Weight/Activation bits: FP4 E2M1 Scale dtype: float8 e4m3fn Group size: 16 Calibration: 512 samples (256 UltraChat + 256 Nemotron CC chat split) MoE calibration: moe calibrate all experts=True Ignore list: Aligned with RedHatAI/Qwen3.5 122B A10B NVFP4 MTP Splicing The upstream heretic quant quantized MTP heads to FP4, which…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy