Qwen3.6 27B uncensored heretic v2 — W8A8 (AutoRound INT8 Dynamic, MTP preserved) INT8 W8A8 dynamic (per token activation) quantization of llmfan46/Qwen3.6 27B uncensored heretic v2 Native MTP Preserved, produced with AutoRound 0.13.1 and exported in the compressed tensors ( auto round:llm compressor ) format for vLLM. Same recipe as the Darwin 28B W8A8 build. Quantization recipe Scheme: INT8 — channel wise INT8 weights, INT8 dynamic per token activations (CUTLASS INT8 path in vLLM). Algorithm: RTN (round to nearest, data free). INT8 + dynamic activations are near lossless. Quantized: language model Linear weights only (DeltaNet in proj / out proj + full attn q/k/v/o proj + MLP gate/up/down ). Preserved in BF16: native MTP module ( mtp. , via ignore layers mtp ), the vision tower ( model.visual. , auto skipped by AutoRound's text module only VLM path), lm head , embed tokens , all norms. Hardware: 2 × RTX 3090. Architecture: Qwen3.5 generation hybrid ( qwen3 5 , 64 layers, 3:1 DeltaNet/full attention, hidden 5120, vocab 248320), multimodal, with a 1 layer native MTP head for speculative decoding. Serving (vLLM, TP2) — recommended: dense (no MTP) MTP note (measured) The native MTP he…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy