Model Overview Description: The NVIDIA DeepSeek R1 0528 FP4 v2 model is the quantized version of the DeepSeek AI's DeepSeek R1 0528 model, which is an auto regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA DeepSeek R1 FP4 model is quantized with TensorRT Model Optimizer. Compared to nvidia/DeepSeek R1 0528 FP4, this checkpoint additionally quantizes the wo module in attention layers. This model is ready for commercial/non commercial use. Third Party Community Consideration This model is not owned or developed by NVIDIA. This model has been developed and built to a third party’s requirements for this application and use case; see link to Non NVIDIA (DeepSeek R1) Model Card. License/Terms of Use: MIT Model Architecture: Architecture Type: Transformers Network Architecture: DeepSeek R1 Input: Input Type(s): Text Input Format(s): String Input Parameters: 1D (One Dimensional): Sequences Other Properties Related to Input: DeepSeek recommends adhering to the following configurations when utilizing the DeepSeek R1 series models, including benchmarking, to achieve the expected performance: \ Set the temperature wit…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy