Model Overview Model Architecture: Llama 3.3 Input: Text Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm : 7.0 PyTorch : 2.8.0 Transformers : 4.53.0 Operating System(s): Linux Inference Engine: vLLM Model Optimizer: AMD Quark (V0.9) Weight quantization: OCP MXFP4, Static Activation quantization: OCP MXFP4, Dynamic KV cache quantization: OCP FP8, Static Calibration Dataset: Pile This model was built with Meta Llama by applying AMD Quark for MXFP4 quantization. Model Quantization This model was obtained by quantizing Llama 3.3 70B Instruct's weights and activations to MXFP4 and KV caches to FP8, using AutoSmoothQuant algorithm in AMD Quark. Quantization scripts: Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend. Evaluation The model was evaluated on MMLU, GSM8K COT, ARC Challenge and IFEVAL. Evaluation was conducted using the framework lm evaluation harness and the vLLM engine. Accuracy Benchmark Llama 3.3 70B Instruct Llama 3.3 70B Instruct MXFP4(this model) Recovery MMLU (5 shot) 83.29 80.99 97.24% GSM8K COT (8 shot, strict match) 93.18 92.12 98.86% ARC Challenge (0 shot) 94.25 93.05 98.73% IFEVAL (0 shot, (inst level str…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy