Qwen3.5 122B A10B FP8 [!Note] This repository contains FP8 quantized model weights and configuration files for the post trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fine grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model. Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. Qwen3.5 Highlights Qwen3.5 features the following enhancement: Unified Vision Language Foundation : Early fusion training on multimodal tokens achieves cross generational parity with Qwen3 and outperforms Qwen3 VL models across reasoning, coding, agents, and visual understanding benchmarks. Efficient Hybrid Architecture : Gated Delta Networks combined with sparse Mixtur…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy