gpt oss 20b WFP8 AFP8 KVFP8 Introduction This model was quantized from openai/gpt oss 20b using AMD Quark with calibration samples from the Pile dataset. Quantization schemes Quantized Layers : All linear (both attention linear and MoE linear) layers excluding lm head Weight : quantized using FP8 symmetric per tensor scheme Activation : quantized using FP8 symmetric per tensor scheme KV Cache : FP8 symmetric per tensor Quantization script Deployment This model supports deployment through the vLLM backend. Please ensure PR 29008, PR 31962 have been correctly applied. Evaluation This model is evaluated on gpqa diamond generative n shot and gsm8k platinum tasks using the lm evaluation harness framework with vllm backend. Evaluation scores Model name Weight Activation KV cache Exclude gpqa diamond generative n shot (5) gsm8k platinum TP1 TP2 TP4 TP8 TP1 TP2 TP4 TP8 openai/gpt oss 20b MXFP4 BF16 BF16 0.5606 0.5303 0.5657 0.5606 0.9016 0.9024 0.9032 0.8966 amd/gpt oss 20b WFP8 AFP8 KVFP8 FP8 FP8 FP8 bias, lm head 0.5505 0.5556 0.5253 0.5253 0.9024 0.9107 0.9024 0.8983 Disclaimer This model is intentionally quantized for the vLLM's CI test usage (tests/models/quantization/test gpt oss.py)…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy