Model Overview Model Architecture: DeepSeek R1 0528 Input: Text Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm : 7.0 PyTorch : 2.8.0 Transformers : 5.0.0 Operating System(s): Linux Inference Engine: SGLang/vLLM Model Optimizer: AMD Quark (V0.11) Base model: Weight quantization: self attn Perchannel, FP8E4M3, Static; MOE OCP MXFP4, Static Activation quantization: self attn Pertoken, FP8E4M3, Dynamic; MOE OCP MXFP4, Dynamic Mtp: Weight quantization: self attn Perchannel, FP8E4M3, Static; MOE OCP MXFP4, Static Activation quantization: self attn Pertoken, FP8E4M3, Dynamic; MOE OCP MXFP4, Dynamic Calibration Dataset: Pile This model was built with deepseek ai DeepSeek R1 0528 model by applying AMD Quark for quantization. Model Quantization The model was quantized from deepseek ai/DeepSeek R1 0528 using AMD Quark. Both weights and activations were quantized. Preprocessing requirement: Before executing the quantization script below, the original FP8 model must first be dequantized to BFloat16. You can either perform the dequantization manually using this conversion script, or use the pre converted BFloat16 model available at amd/DeepSeek R1 0528 BF16. Quantization…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy