Dataset Card for "wiki dpr" Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Repository: https://github.com/facebookresearch/DPR Paper: https://arxiv.org/abs/2004.04906 Point of Contact: More Information Needed Dataset Summary This is the wikipedia split used to evaluate the Dense Passage Retrieval (DPR) model. It contains 21M passages from wikipedia along with their DPR embeddings. The wikipedia articles were split into multiple, disjoint text blocks of 100 words as passages. The wikipedia dump is the one from Dec. 20, 2018. There are two types of DPR embeddings based on two different models: nq : the model is trained on the Natural Questions dataset multiset : the model is trained on multiple datasets Additionally, a FAISS index can be created from the embeddings: ex…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy