Gemma 4 31B Dense AWQ 4 bit In house AWQ 4 bit calibration of google/gemma 4 31b it, end to end from the upstream BF16 base. Thinking + vision aware calibration via balanced thinking vision corpus (40% AM Thinking v1 Distilled / 30% LLaVA Instruct / 15% NuminaMath / 15% UltraChat). Replaces the older mattbucci/gemma 4 31B it AutoRound AWQ which was a repack of Intel's AutoRound GPTQ output (50.4% negative scales). This ship is fully in house: standard AWQ scales, thinking traces preserved, vision tower kept BF16. Model Details Base model google/gemma 4 31b it Architecture Dense with sliding window attention (50 SWA + 10 full attention layers) Parameters 31B Layers 60 Quantization AWQ 4 bit, group size=128 Calibration 512 samples × 1024 tokens, balanced thinking vision recipe (text only — vision tower BF16) Scale audit 0 / 410 quantized tensors flagged (clean) Capability Validation (R9700 / SGLang v0.5.11) Probe Result Notes : : basic ("What is the capital of France?") ✅ clean 'paris', finish=stop thinking ✅ 460 tok reasoning, terminated cleanly vision (red circle on white) ⚠ crashes see Known Limitations Known Limitations Vision: BROKEN on RDNA4. The model generates a coherent visi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy