Gemma 4 26B A4B IT Heretic (FP8 Static) This repository contains an offline statically quantized FP8 version of the coder3101/gemma 4 26B A4B it heretic model. The quantization procedure was designed to maximize inference throughput on NVIDIA hardware equipped with FP8 tensor cores (e.g., Hopper, Blackwell architectures) while preserving conversational accuracy. Quantization Methodology The model was quantized utilizing the AutoFP8 library. To ensure optimal runtime performance without the overhead of dynamic activation scaling, a strict static activation scheme was employed. Quantization Format : FP8 (W8A8) Activation Scheme : Static Calibration Dataset : 512 samples from the mgoin/ultrachat 2k dataset. Exclusions for Multimodality & Architecture Stability : The Gemma 4 architecture utilizes separate pathways for vision (the Vision Tower) and text generation. AutoFP8 implicitly attempts to quantize all linear layers indiscriminately. To ensure the model remains fully multimodal and compatible with vLLM, a surgical restoration process was applied post quantization: 1. The config.json was patched to explicitly add the vision tower and embed vision modules, alongside the MoE router.p…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy