Gemma 4 12B it AEON Abliterated — K=4 Biprojection (FP8) The near lossless 8 bit FP8 quantization of our K=4 biprojection abliteration of google/gemma 4 12B it . Matches BF16 capability (MMLU, HumanEval, IFEval all within noise) at ~half the size and 1.6× the throughput. Loads in vLLM with quantization modelopt . ~13 GB. This is the recommended variant when quality matters. For maximum speed/smallest size (with a measured reasoning trade off) see the NVFP4 sibling. Refusal behavior has been removed; the model responds to a wide range of prompts the base would decline. Operator side safety is your responsibility — see the arbitration clause at the bottom. 🚀 QuickStart Complete copy paste recipe for DGX Spark / Blackwell — pull the container, download the model, serve. This is a plain decode variant (no speculative drafter — MTP is net neutral on the GB10), so there is no drafter to pull and no speculative config . Unified memory gpu util: On the DGX Spark's unified memory keep gpu memory utilization at 0.6 0.7; above ~0.8 the shared CPU+GPU pool page thrashes. Discrete VRAM GPUs can run higher. Then call it (OpenAI compatible): Plain vLLM (pip, non container) ⚠️ Needs vLLM ≥ 0.23.0…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy