Model Overview Model Architecture: DeepSeek R1 0528 Input: Text Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm : 7.0 PyTorch : 2.8.0 Transformers : 4.53.0 Operating System(s): Linux Inference Engine: SGLang/vLLM Model Optimizer: AMD Quark (V0.10) Weight quantization: OCP MXFP4, Static Activation quantization: OCP MXFP4, Dynamic Calibration Dataset: Pile This model was built with deepseek ai DeepSeek R1 0528 model by applying AMD Quark for MXFP4 quantization. Model Quantization The model was quantized from deepseek ai/DeepSeek R1 0528 using AMD Quark. Both weights and activations were quantized to MXFP4 format. Preprocessing requirement: Before executing the quantization script below, the original FP8 model must first be dequantized to BFloat16. You can either perform the dequantization manually using this conversion script, or use the pre converted BFloat16 model available at amd/DeepSeek R1 0528 BF16. Quantization scripts: Deployment This model can be deployed efficiently using the SGLang and vLLM backends. Evaluation The model was evaluated on AIME24, and GSM8K benchmarks using the lm evaluation harness framework. Accuracy Benchmark DeepSeek R1 0528 MXFP4…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy