📢 Updesh: Synthetic Multilingual Instruction Tuning Dataset for 13 Indic Languages NOTE: This is an initial $\beta$ release. We plan to release subsequent versions of Updesh with expanded coverage and enhanced quality control. Future iterations will include larger datasets, improved filtering pipelines. Updesh is a large scale synthetic dataset designed to advance post training of LLMs for Indic languages. It integrates translated reasoning data and synthesized open domain generative content to support culturally grounded multilingual adaptation of LLMs. Despite the rapid progress in instruction tuned LLMs, most existing datasets focus on English, creating a gap in high quality, culturally grounded resources for Indic languages—resources that are essential for enabling Small Language Models (SLMs) to serve India’s diverse linguistic landscape. Updesh aims to fill this gap by providing rich, multilingual instruction tuning data grounded in Indian languages and contexts. Unlike previous English centric translated datasets, Updesh employs a dual approach of culturally grounded data generation and careful, selective translation, ensuring linguistic nuance and relevance for each langua…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy