Want a bigger model? Download Sarvam 105B! Index 1. Introduction 2. Architecture 3. Benchmarks Knowledge & Coding Reasoning & Math Agentic 4. Inference Hugging Face vLLM SGLang 5. Footnote 6. Citation Introduction Sarvam 30B is an advanced Mixture of Experts (MoE) model with 2.4B non embedding active parameters, designed primarily for practical deployment. It combines strong reasoning, reliable coding ability, and best in class conversational quality across Indian languages. Sarvam 30B is built to run reliably in resource constrained environments and can handle multilingual voice calls while performing tool calls. A major focus during training was the Indian context and languages, resulting in state of the art performance across 22 Indian languages for its model size. Sarvam 30B is open sourced under the Apache License . For more details, see our blog. Architecture The 30B MoE model is designed for throughput and memory efficiency. It uses 19 layers, a dense FFN intermediate size of 8192, moe intermediate size of 1024, top 6 routing, grouped KV heads ( num key value heads=4 ), and an extremely high rope theta ( 8e6 ) for long context stability without RoPE scaling. It has 128 exper…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy