GRM 2.6 Plus W8A8 (INT8 weights + INT8 dynamic activations) — MTP preserved W8A8 INT8 quantization of OrionLLM/GRM 2.6 Plus, produced with llm compressor (SmoothQuant + GPTQ). The native MTP (Multi Token Prediction) head is preserved as BF16 and works at 91 94 % acceptance under vLLM ≥ 0.17 — running ~26 % faster than the BF16 source while taking 35 % less memory. Quick start No additional patches required — both config.json.quantization config.ignore (covers MTP Linear modules) and actorder field are already fixed. Performance (single H200 SXM, vLLM 0.17.1, temperature=0) Workload Concurrency Throughput TPOT p50 MTP Accept rate Code generation 1 131 tok/s 7.6 ms 90.8 % Code generation 8 859 tok/s 8.5 ms 91.1 % JSON structured output 1 132 tok/s 7.6 ms 94.2 % JSON structured output 8 942 tok/s 8.3 ms 93.7 % For reference, the BF16 source with the same MTP recipe reaches 102 tok/s single stream and 749 tok/s at concurrency 8 on W2 — i.e. this W8A8 model is 26 29 % faster while using ~18 GB less weight memory. Architecture preserved Component Status Language model Linear (q/k/v/o, MLP) on the 16 full attention layers INT8 (W8A8 channelwise weight + dynamic per token activation) MLP o…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy