Want a smaller model? Download Sarvam 30B! Index 1. Introduction 2. Architecture 3. Benchmarks Knowledge & Coding Reasoning & Math Agentic 4. Inference Hugging Face vLLM SGLang 5. Footnote 6. Citation Introduction Sarvam 105B is an advanced Mixture of Experts (MoE) model with 10.3B active parameters, designed for superior performance across a wide range of complex tasks. It is highly optimized for complex reasoning, with particular strength in agentic tasks, mathematics, and coding. Sarvam 105B is a top tier performer, consistently matching or surpassing several major closed source models and staying within a narrow margin of frontier models across diverse reasoning and agentic benchmarks. It demonstrates exceptional agentic and reasoning capabilities in real world applications such as web search and technical troubleshooting. A major focus during training was the Indian context and languages, resulting in state of the art performance across 22 Indian languages for its model size. Sarvam 105B is open sourced under the Apache License . For more details, see our blog. Architecture The 105B model adopts an MLA style attention stack with decoupled QK head dimensions ( q head dim=192 sp…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy