Gemma-4-E4B-it-NVFP4
NVFP4 quantized version of google/gemma-4-E4B-it for vLLM.
Quantization Profile
- Text backbone:
NVFP4 lm_head: higher precision- Vision tower and vision embeddings: higher precision
- Audio tower and audio embeddings: higher precision
- KV cache:
FP8
Usage
vllm serve Neural-ICE/Gemma-4-E4B-it-NVFP4 \
--quantization modelopt \
--gpu-memory-utilization 0.90
Official Gemma 4 vLLM recipe:
https://docs.vllm.ai/projects/recipes/en/latest/Google/Gemma4.html
Text Generation
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Neural-ICE/Gemma-4-E4B-it-NVFP4",
"messages": [
{"role": "user", "content": "Explain quantum entanglement in simple terms."}
],
"max_tokens": 512
}'