Model Overview Description: The NVIDIA GLM 5 NVFP4 model is the quantized version of ZAI’s GLM 5 model, which is an auto regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA GLM 5 NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non commercial use. Third Party Community Consideration This model is not owned or developed by NVIDIA. This model has been developed and built to a third party’s requirements for this application and use case; see link to Non NVIDIA (GLM 5) Model Card from ZAI. References Nvidia Model Optimizer: https://github.com/NVIDIA/Model Optimizer License/Terms of Use: MIT License Deployment Geography: Global Use Case: Developers looking to take off the shelf, pre quantized models for deployment in AI Agent systems, chatbots, RAG systems, and other AI powered applications. Release Date: Huggingface 03/16/2026 via https://huggingface.co/nvidia/GLM 5 NVFP4 Model Architecture: Architecture Type: Transformers Network Architecture: GLM 5 Number of Model Parameters: 744B in total and 40B activated Input: Input Type(s): Text Input Format(s): String Input Parameters: On…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy