Model Overview Description: The NVIDIA GLM 5.2 NVFP4 model is the quantized version of ZAI’s GLM 5.2 model, which is an auto regressive language model that uses an optimized transformer architecture. GLM 5.2 is a Mixture of Experts (MoE) model for reasoning and coding that uses sparse attention (with an IndexShare indexer) to support a long context. For more information, please check here. The NVIDIA GLM 5.2 NVFP4 model is quantized with Model Optimizer. This model is ready for commercial or non commercial use. License/Terms of Use: GOVERNING TERMS: Use of the model is governed by the MIT License, same as the base model. Deployment Geography: Global Use Case: Developers looking to take off the shelf, pre quantized models for deployment in AI Agent systems, chatbots, RAG systems, and other AI powered applications. Release Date: Hugging Face 06/25/2026 via https://huggingface.co/nvidia/GLM 5.2 NVFP4 References Nvidia Model Optimizer: https://github.com/NVIDIA/Model Optimizer Model Architecture: Architecture Type: Transformers Network Architecture: GLM 5.2 ( GlmMoeDsaForCausalLM ) Number of Model Parameters: 753B in total and 40B activated Input: Input Type(s): Text Input Format(s): S…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy