amd/gpt oss 20b MoE Quant W MXFP4 A FP8 KV FP8 Introduction This model was quantized from openai/gpt oss 20b using AMD Quark with calibration samples from the Pile dataset. Quantization schemes Quantized Layers : All linear layers excluding self attn , router , lm head , i.e., only MoE MLP linear layers are quantized. Weight : quantized using OCP Microscaling (MX) FP4 scheme Activation : quantized using FP8 symmetric per tensor scheme KV Cache : FP8 symmetric per tensor Quantization script Deployment This model supports deployment through the vLLM backend. Please ensure PR 29008 has been correctly applied. Evaluation This model is evaluated on gpqa diamond generative n shot and gsm8k platinum tasks using the lm evaluation harness framework with vllm backend. Model name Weight Activation KV cache Exclude gpqa diamond generative n shot (5) gsm8k platinum TP1 TP2 TP4 TP8 TP1 TP2 TP4 TP8 openai/gpt oss 20b MXFP4 BF16 BF16 0.5606 0.5303 0.5657 0.5606 0.9016 0.9024 0.9032 0.8966 amd/gpt oss 20b WMXFP4 AFP8 KVFP8 (this one) MXFP4 FP8 FP8 self attn router lm head 0.5303 0.5556 0.5152 0.5404 0.8999 0.8900 0.8958 0.9098 Disclaimer This model is intentionally quantized for the vLLM's CI test…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy