Dynamic 4 bit quantization of black forest labs/FLUX.2 klein 4B using SDNQ. This model uses per layer fine grained quantization. What dtype to use for a layer is selected dynamically by trial and error until the std normalized mse loss is lower than the selected threshold. Minimum allowed dtype is set to uint4 and std normalized mse loss threshold is set to 1e 2. This created a mixed precision model with uint4 and int5 dtypes. SVD quantization is disabled. Usage: Original BF16 vs SDNQ quantization comparison: Quantization Model Size Visualization Original BF16 7.8 GB SDNQ 4 Bit 2.5 GB
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy