DeepSeek V4 Flash NVFP4 FP8 Model Optimizations This model was obtained by using the following branch with LLM Compressor: https://github.com/vllm project/llm compressor/pull/2647 Deployment This model was deployed using the following branch with vLLM: https://github.com/vllm project/vllm/pull/41276 Evaluation This model has a noticably lower accuracy recovery than the base model due to the base model being released in a quantized format and differences between mxfp4 and nvfp4. More advanced techniques such as GPTQ can be used to increase accuracy recovery beyond this model's current state. For more details on how this model was created and run in LLM Compressor, please contact Kyle Sayers on the vLLM Slack: https://communityinviter.com/apps/vllm dev/join vllm developers slack
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy