Model Overview Model Architecture: gpt oss 120b Input: Text Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm : 7.0 Operating System(s): Linux Inference Engine: vLLM Model Optimizer: AMD Quark Weight quantization: OCP MXFP4, Static Activation quantization: OCP MXFP4, Dynamic Calibration Dataset: Pile This model was built with gpt oss 120b model by applying AMD Quark for MXFP4 quantization. Model Quantization The model was quantized from openai/gpt oss 120b using AMD Quark. The weights are quantized MXFP4 and activations were quantized to FP8. Quantization scripts: Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend. Evaluation The model was evaluated on AIME25 and GPQA Diamond benchmarks with low reasoning effort. Accuracy Benchmark gpt oss 120b gpt oss120b w mxfp4 a fp8(this model) Recovery AIME25 65.25 67.12 102.87% GPQA 51.67 53.42 103.39% Reproduction The results of AIME25 and GPQA Diamond were obtained using gpt oss.evals with low effort setting, and vLLM docker rocm/vllm private:rocm7.0 ubuntu 22.04 vllm 0.10.1 instinct gptoss wmxfp4 afp8 20251030 . Launching server Evaluating model in a new terminal License Modificatio…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy