SCIDOCS An MTEB dataset Massive Text Embedding Benchmark SciDocs, a new evaluation benchmark consisting of seven document level tasks ranging from citation prediction, to document classification and recommendation. Task category t2t Domains Academic, Written, Non fiction Reference https://allenai.org/data/scidocs How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: To learn more about how to run models on mteb task check out the GitHub repitory. Citation If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as a part of the MMTEB Contribution. Dataset Statistics Dataset Statistics The following code contains the descriptive statistics from the task. These can also be obtained using: This dataset card was automatically generated using MTEB
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy