Qwen3.6 27B AEON Ultimate Uncensored NVFP4 NVFP4 hardware quantized release of AEON 7/Qwen3.6 27B AEON Ultimate Uncensored . Same model, same 0/50 refusal rate, same preserved capabilities — now compressed from 51 GB BF16 to 26 GB NVFP4 for native FP4 hardware throughput on DGX Spark / GB10 / Blackwell sm 121a. This release is multimodal preserved (vision tower stays BF16 — text + image inference fully functional) and hybrid attention preserved (the 48 linear attention / GatedDeltaNet layers stay BF16; FP4 only applies to the 16 full attention layers' output projections + all MLPs). What Changed vs BF16 Aspect BF16 (source) NVFP4 (this release) Disk size 51 GB 26 GB (49% reduction) Refusal rate 0/50 inherited — to be verified post deploy Multimodal preserved preserved (vision BF16, no degradation) Hybrid SSM repaired + intact intact (linear attn BF16 preserved) Hardware target A100 / H100 / RTX PRO 6000 BF16 DGX Spark (GB10), B100/B200, RTX PRO 6000 Blackwell with native FP4 throughput KL vs BF16 source n/a expected ≤0.001 (typical for this recipe class) The NVFP4 quantization scheme is NVIDIA mandated: E2M1 element format, block size=16, FP8 E4M3 per block scales, FP32 per tensor…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy