MIRACLRetrievalHardNegatives An MTEB dataset Massive Text Embedding Benchmark MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5 multilingual large and e5 mistral instruct. Task category t2t Domains Encyclopaedic, Written Reference http://miracl.ai/ Source datasets: mteb/MIRACLRetrieval mteb/miracl hard negatives How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repository. Citation If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as a part of the MMTEB Contribution. Dataset Statistics Dataset Statistics The following code contains the descriptive statistics from the task. These can also be obtained using: This dataset card was automatically generated using MTEB
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy