Multilingual Embeddings for Wikipedia in 300+ Languages This dataset contains the wikimedia/wikipedia dataset dump from 2023 11 01 from Wikipedia in all 300+ languages. The individual articles have been chunked and embedded with the state of the art multilingual Cohere Embed V3 embedding model. This enables an easy way to semantically search across all of Wikipedia or to use it as a knowledge source for your RAG application. In total is it close to 250M paragraphs / embeddings. You can also use the model to perform cross lingual search: Enter your search query in any language and get the most relevant results back. Loading the dataset Loading the document embeddings You can either load the dataset like this: Or you can also stream it without downloading it before: Note, depending on the language, the download can be quite large. Search A full search example (on the first 1,000 paragraphs): Overview The following table contains all language codes together with the total numbers of passages. Language Docs : : en 41,488,110 de 20,772,081 fr 17,813,768 ru 13,734,543 es 12,905,284 it 10,462,162 ceb 9,818,657 uk 6,901,192 ja 6,626,537 nl 6,101,353 pl 5,973,650 pt 5,637,930 sv 4,911,480 c…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy