RnJ 1 Instruct FP8 Model Description This is an FP8 quantized version of EssentialAI/RnJ 1 Instruct, created using llmcompressor (Neural Magic). Key Benefits: ~50% smaller model size (8GB vs 16GB) Native FP8 inference on Ada Lovelace, Hopper, and Blackwell GPUs Single consumer GPU deployment on 12GB+ cards (RTX 3060, RTX 4070, etc.) Native vLLM and SGLang support Minimal quality loss with FP8 dynamic quantization Key Features RnJ 1 Instruct is a Gemma3 based instruction following model with: Strong Math Performance : GSM8K 87.19% (5 shot) Multi Domain Knowledge : MMLU Pro 44.45% Efficient Architecture : Only 8B parameters, fast inference 32K Context : Extended context window for documents Quantization Details Property Value Quantization Method FP8 Dynamic (W8A8) Weights Precision FP8 E4M3 (8 bit) Activations Precision FP8 E4M3 (8 bit, dynamic) Ignored Layers lm head (kept in BF16) Quantization Tool llmcompressor 0.12.2 Original Model Size ~16GB Quantized Model Size ~8GB Quantization Recipe Quick Start with Docker The easiest way to run this model. No setup required just Docker with NVIDIA runtime. Docker Compose (Recommended) Docker Run Test the API Usage vLLM (Recommended) SGLang…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy