The FLUX.2 [klein] model family are our fastest image models to date. FLUX.2 [klein] unifies generation and editing in a single compact architecture, delivering state of the art quality with end to end inference in as low as under a second . Built for applications that require real time image generation without sacrificing quality, and runs on consumer hardware, with as little as 13GB VRAM. FLUX.2 [klein] 4B is a 4 billion parameter rectified flow transformer capable of generating images from text descriptions and supports multi reference editing capabilities. Fully open under Apache 2.0. Our most accessible model runs on consumer GPUs like the RTX 3090/4070. Compact but capable: supports text to image, image editing, and multi reference at quality that punches above its size. Built for local development, edge deployment, and production use. For more information, please read our blog post. Quantization Details 🔧 This model is a quantized version optimized for efficient inference: Transformer : Quantized using TorchAo int8 (int8wo) quantization, significantly reducing model size while maintaining generation quality. Text Encoder : Replaced with unsloth/Qwen3 4B unsloth bnb 4bit , a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy