Sangraha Sangraha is the largest high quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations. More information : For detailed information on the curation and cleaning process of Sangraha, please checkout our paper on Arxiv; Check out the scraping and cleaning pipelines used to curate Sangraha on GitHub; Getting Started For downloading the entire Sangraha: For downloading a subset (Verified/Unverified) of Sangraha: For downloading one language from a subset of Sangraha: Background Sangraha contains three broad components: Sangraha Verified : Containing scraped data from "human verified" Websites, OCR extracted data from high quality Indic language PDFs, transcribed data from various Indic language videos, podcasts, movies, courses, etc. Sangraha Unverfied : High quality Indic language data extracted from existing multilingual corpora employing perplexity filtering using n gram language models trained on Sangraha Verified. Sangraha Synthetic : WikiMedia English translated to 14 Indic languages and further "romanised" from 14 languages by translitera…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy