Param 2 17B BharatGen presents Param 2 17B MoE A2.4B , a large scale Mixture of Experts (MoE) language model designed to deliver high model capacity while retaining the inference efficiency of a much smaller dense model. It uses a Hybrid MoE architecture with 17B total parameters , while activating only 2.4B parameters per token . The model is pretrained from scratch, with a strong emphasis on linguistic diversity , cultural grounding , and multilingual representation , particularly for Indian languages. It is released as an early post training checkpoint with advanced capabilities including reasoning, tool calling, mathematics, and code generation, making it suitable for diverse downstream applications and further fine tuning. 🌟 Key Highlights 17B parameter Mixture of Experts (MoE) language model Multilingual : English, Hindi + 21 Indian languages Trained on ~22 trillion tokens across two pretraining phases Uses 64 specialized experts , dynamically activated per token Supports long context understanding (up to 4096 tokens) Efficient inference : Only 2.4B active parameters per token Advanced Capabilities: Thinking & Reasoning, Tool Calling, Mathematics, Code Generation Designed fo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy