Qwen3.6 35B A3B Hybrid INT4 FP8 MTP Hybrid quantization of Qwen/Qwen3.6 35B A3B optimized for single GPU deployment on the NVIDIA DGX Spark / ASUS GX10 (GB10, 128 GB unified memory) . MoE expert FFN layers → INT4 (Intel AutoRound group size=128 ) Attention, shared experts, LM head, embeddings → FP8 (E4M3 blockwise 128×128, calibrated by Qwen) MTP (Multi Token Prediction) drafter weights included for speculative decoding The result is a ~20 GB checkpoint (vs. ~70 GB BF16, ~35 GB FP8) that runs at ~92–103 tok/s single request and ~164 tok/s aggregate at 16× concurrency on a single GB10, while keeping headroom for 256k context KV cache . Headline numbers Measured on ASUS GX10 / DGX Spark (GB10, 128 GB unified memory) with the patched vLLM runtime described below, single request unless stated otherwise. Metric Value Single request peak (short context) ~103 tok/s Concurrent peak total (16 parallel) ~164 tok/s Prompt processing throughput (pp=4096) ~5500 tok/s TTFT @ 4k prompt ~750 ms Disk size (safetensors) ~20 GB Max context 262144 For reference, phuongncn's README reports for the same hardware: Build Tokens/sec Ollama (Q4 GGUF) ~30 tok/s llama.cpp (manual SM121) ~49 tok/s Qwen3.5 35B…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy