Model Overview Model Architecture: DeepSeek R1 Input: Text Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm : 7.0 PyTorch : 2.8.0 Transformers : 4.53.0 Operating System(s): Linux Inference Engine: SGLang Model Optimizer: AMD Quark (V0.10) Weight quantization: OCP MXFP4, Static Activation quantization: OCP MXFP4, Dynamic KV cache : OCP FP8, Static Calibration Dataset: Pile This model was built with deepseek ai DeepSeek R1 model by applying AMD Quark for MXFP4 quantization. Model Quantization The model was quantized from deepseek ai/DeepSeek R1 using AMD Quark. Both weights and activations were quantized to MXFP4 format. Preprocessing requirement: Before executing the quantization script below, the original FP8 model must first be dequantized to BFloat16. You can either perform the dequantization manually using this conversion script, or use the pre converted BFloat16 model available at unsloth/DeepSeek R1 BF16. Quantization scripts: Deployment Use with SGLang This model can be deployed efficiently using the SGLang backend. Evaluation The model was evaluated using SGLang and lm evaluation harness frameworks. Accuracy Benchmark DeepSeek R1 DeepSeek R1 MXFP4(this…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy