Step3 VL 10B AWQ Base model: stepfun ai/Step3 VL 10B I added a small injection at the end of the original chat template.jinja to support a GLM like switch: you can try to disable the “thinking/reasoning” mode in your chat completion request via "chat template kwargs": {"enable thinking": False} . Note that the model was not specifically trained with this switch, so it may not reliably follow it in all cases. 【Dependencies / Installation】 As of 2026 01 22 , make sure your system has cuda12.8 installed. Then, create a fresh Python environment (e.g. python3.12 venv) and run: 【vLLM Startup Command】 【Logs】 【Model Files】 File Size Last Updated 8 GiB 2026 01 22 【Model Download】 【Overview】 STEP3 VL 10B 📢 News & Updates 🚀 Online Demo : Explore Step3 VL 10B on ModelScope Spaces ! 📢 [Notice] vLLM Support: vLLM integration is now officially supported! (PR 32329) ✅ [Fixed] HF Inference: Resolved the eos token id misconfiguration in config.json that caused infinite generation loops. (PR f55c07e) ✅ [Fixing] Metric Correction: We sincerely apologize for inaccuracies in the Qwen3VL 8B benchmarks (e.g., AIME, HMMT, LCB). The errors were caused by an incorrect max tokens setting (mistakenly set to…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy