Gemma 4 E2B it NVFP4A16 NVFP4 quantization of google/gemma 4 E2B it — Google's Gemma 4 E2B instruct (chat optimized) with Per Layer Embeddings (PLE). 2B effective parameters, multimodal (text + image + audio), 128K context. W4A16 — language model weights in FP4, activations in FP16. Vision and audio towers stay BF16. See also Gemma 4 E2B it NVFP4 for the full W4A4 variant. Key Specs Original (BF16) NVFP4 (this) Size on disk ~10 GB ~7.5 GB Compression — ~1.3x (text layers 3x, vision/audio stay BF16) Effective parameters 2B 2B Architecture Dense + PLE (Per Layer Embeddings) same Context window 128K tokens 128K tokens Modalities Text, Image, Audio Text, Image, Audio What is PLE? Unlike the Gemma 4 26B which uses Mixture of Experts (MoE), the E2B uses Per Layer Embeddings — a learned per layer specialization mechanism. Each of the 35 decoder layers gets its own 256 dimensional signal derived from both the token identity (via a second embedding table) and the evolving hidden representation. This is a continuous alternative to discrete MoE routing — no expert selection, no sparsity, just dense computation with layer specific conditioning. Speed Tested on DGX Spark (GB10 Blackwell, SM 12.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy