Dataset Card for "ms marco" Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: https://microsoft.github.io/msmarco/ Repository: More Information Needed Paper: More Information Needed Point of Contact: More Information Needed Size of downloaded dataset files: 1.55 GB Size of the generated dataset: 4.72 GB Total amount of disk used: 6.28 GB Dataset Summary Starting with a paper released at NIPS 2016, MS MARCO is a collection of datasets focused on deep learning in search. The first dataset was a question answering dataset featuring 100,000 real Bing questions and a human generated answer. Since then we released a 1,000,000 question dataset, a natural langauge generation dataset, a passage ranking dataset, keyphrase extraction dataset, crawling dataset, and a conv…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy