Qwythos 9B v2 NVFP4 NVFP4 (W4A4) build of empero ai/Qwythos 9B v2 for vLLM on NVIDIA Blackwell — the native 4 bit serving path nobody ships for this checkpoint. Empero already publishes the MTP GGUF for llama.cpp; this is the piece that was missing: a compressed tensors NVFP4 artifact that runs the model's GEMMs directly on the Blackwell FP4 tensor cores under vLLM. Quantized by protoLabs. Base model, its capabilities, and its uncensored research posture are Empero AI's — see their card. What this is NVFP4 W4A4 on the 128 transformer linears (attention + MLP). The hybrid DeltaNet ( linear attn ) layers, the vision tower, lm head , and the MTP head are kept BF16 — DeltaNet corrupts under 4 bit, and the vision path stays lossless. 11.2 GB on disk (vs 19.3 GB BF16 source; the BF16 preserved vision tower is most of the remainder). MTP sidecar included ( model mtp.safetensors ) — see the status note below. Inherits the base wholesale: 1M token context (YaRN), Qwen3.5 multimodal stack, FTPO looping fix (greedy loop rate 6.7%→0%), uncensored for research/red team/bio chem clinical work. Speed Single RTX PRO 6000 Blackwell (sm120), vLLM 0.24.0, NVFP4 linear backend marlin , BF16 KV: Regime…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy