Qwen3.5 122B A10B int4 AutoRound EC Extended Calibration (EC) INT4 AutoRound quantization of Qwen/Qwen3.5 122B A10B, a 122B MoE (10B active) multimodal model. Drop in replacement for Intel/Qwen3.5 122B A10B int4 AutoRound with wider calibration settings for improved quality on long context and reasoning heavy workloads. Calibration — Extended vs Intel default Intel (v0.12.0) EC (this model) iters 200 400 nsamples 128 256 seqlen 512 4096 batch size 1 8 (default) grad accum 8 1 (default) ignore layers shared expert shared expert bits / group 4 / 128 4 / 128 Effective calibration batch size is 8 in both cases (1×8 vs 8×1) — mathematically equivalent signal per optimizer step, just different memory/latency profile during quantization. Environment Component Version auto round 0.12.2 transformers 5.5.3 torch 2.11.0 safetensors 0.7.0 huggingface hub 1.10.1 Hardware RunPod H200 SXM (1x) Wall time ~15 hrs Files Path What model 000{01..13} of 00013. Quantized language model shards (INT4 GPTQ, w4g128) model visual.safetensors Visual encoder (BF16, base model passthrough) model extra tensors.safetensors MTP (multi token prediction) weights (BF16 passthrough) config.json Multimodal config with…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy