SmallThinker 4BA0.6B Instruct GGUF GGUF models with .gguf suffix can used with llama.cpp framework. GGUF models with .powerinfer.gguf suffix are integrated with fused sparse FFN operators and sparse LM head operators. These models are only compatible to powerinfer framework. Introduction   🤗 Hugging Face      🤖 ModelScope       📑 Technical Report    SmallThinker is a family of on device native Mixture of Experts (MoE) language models specially designed for local deployment, co developed by the IPADS and School of AI at Shanghai Jiao Tong University and Zenergize AI . Designed from the ground up for resource constrained environments, SmallThinker brings powerful, private, and low latency AI directly to your personal devices, without relying on the cloud. Performance Note: The model is trained mainly on English. Model MMLU GPQA diamond GSM8K MATH 500 IFEVAL LIVEBENCH HUMANEVAL Average : : : : : : : : : : : : : : : : : SmallThinker 4BA0.6B Instruct 66.11 31.31 80.02 60.60 69.69 42.20 82.32 61.75 Qwen3 0.6B 43.31 26.77 62.85 45.6 58.41 23.1 31.71 41.67 Qwen3 1.7B 64.19 27.78 81.88 63.6 69.50 35.60 61.59 57.73 Gemma3nE2b it 63.04 20.2 8…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy