NVIDIA Nemotron 3 Ultra 550B A55B FP8 block Model Overview Model Architecture: NemotronHForCausalLM Input: Text Output: Text Total Parameters: 550B Active Parameters: 55B Model Optimizations: Weight quantization: FP8 (per block) Activation quantization: FP8 (dynamic per token) Intended Use Cases: Reasoning and complex problem solving. Mathematics and science. Code generation. Instruction following. Out of scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Release Date: 06/04/2025 Version: 1.0 Model Developers: Red Hat Quantized version of nvidia/NVIDIA Nemotron 3 Ultra 550B A55B BF16. Model Optimizations This model was obtained by quantizing the weights and activations of nvidia/NVIDIA Nemotron 3 Ultra 550B A55B BF16 to FP8 data type. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformer blocks are quantized. Weights are quantized with FP8 per block quantization, while activations are quantized with FP8 dynamic per token quantization. The llm compressor library is…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy