EXPERIMENTAL feedback welcome AWQ quantization: done by stelterlab in INT4 GEMM with llm compressor (https://github.com/vllm project/llm compressor v0.9.0.1) from the vllm project. See recipe.yaml. Did some basic tests against vLLM v0.15.0 (official docker image). Original Weights by NVIDIA. Original Model Card follows: NVIDIA Nemotron 3 Nano 30B A3B AWQ (4 bit) Model Overview Model Developer: NVIDIA Corporation Model Dates: September 2025 \ December 2025 Data Freshness: The post training data has a cutoff date of November 28, 2025\. The pre training data has a cutoff date of June 25, 2025\. Description Nemotron Nano 3 30B A3B FP8 is a quantized version of Nemotron Nano 3 30B A3B and is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy