Grok 1 W4A8KV8 Introduction This model was created by applying Quark with calibration samples from Pile dataset. Quantization Stragegy Quantized Layers : All linear layers excluding "lm head", " .gate" Weight : FP8 symmetric per tensor, additionally, INT4 symmetric per channel for MoE linear Activation : FP8 symmetric per tensor KV Cache : FP8 symmetric per tensor INT4 Packing Every eight int4 values are packed into a single int32 integeter following the sequence defined by order map = [0, 2, 4, 6, 1, 3, 5, 7] . Quick Start Follow Quantizing Sharded Grok 1 with Quark for SGLang to produced the quantized model using Quark. Deployment Quark has its own export format and allows FP8 quantized models to be efficiently deployed using the SGLang backend. Evaluation Evaluation scores Benchmark grok 1 grok 1 W4A8KV8(this model) gsm8k 0.821 0.817 License Modifications copyright(c) 2024 Advanced Micro Devices,Inc. All rights reserved.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy