Qwen3.6 35B A3B Uncensored HauhauCS Aggressive NVFP4 NVFP4 quantized version of HauhauCS/Qwen3.6 35B A3B Uncensored HauhauCS Aggressive. Conservative profile : linear attn (30 DeltaNet/Mamba layers) and MTP kept in bf16 for best quality. Follows AEON 7/RedHatAI approach. Spec Value Base model HauhauCS/Qwen3.6 35B A3B Uncensored HauhauCS Aggressive (Q8 K P GGUF) Original model Qwen/Qwen3.6 35B A3B Architecture Qwen3.5 MoE — 35B total, 3B active, 256 experts (8 routed + 1 shared) Quantization NVFP4 W4A4 (conservative: linear attn + MTP in bf16) Format compressed tensors (native vLLM support) Size ~22 GB Max context (text only) 131K+ on RTX 5090 Requires NVIDIA Blackwell GPU (SM 120) Quantization Recipe Calibration: HuggingFaceH4/ultrachat 200k, 128 samples × 1024 tokens MTP tensors copied from Qwen/Qwen3.6 35B A3B (not present in GGUF) Deployment (vLLM) Vision + text smoke tested on RTX 5090 This repository has been smoke tested locally on an RTX 5090 with vllm/vllm openai:v0.21.0 cu130 local , compressed tensors , NVFP4 Marlin GEMM, FP8 KV cache, and a real image chat.completions request. For short non thinking answers, pass chat template kwargs at the top level of the OpenAI compat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy