NanoMTEB Scandinavian This dataset is a Nano style retrieval dataset for HAKARI bench. NanoMTEB Scandinavian is a compact retrieval benchmark for Scandinavian language MTEB style task families. It includes Danish, Norwegian, and Swedish retrieval tasks spanning fact verification, question answering, news, encyclopedic content, FAQ retrieval, and social media retrieval. Usage Data Layout This dataset uses six Hugging Face Datasets configs: corpus : documents with id and text queries : queries with id and text qrels : positive relevance labels with query id and corpus id bm25 : BM25 candidate lists with query id and corpus ids harrier oss v1 270m : dense candidate lists from microsoft/harrier oss v1 270m reranking hybrid : RRF candidate lists built from bm25 and harrier oss v1 270m Each config has the same Nano split names. Candidate Construction bm25 : local BM25 top 500 with automatic language aware tokenization. The resolved tokenizer is shown in the Candidate Quality table, for example wordseg@ja . harrier oss v1 270m : dense top 500 from microsoft/harrier oss v1 270m . In tables this is shown as Dense ; Dense means microsoft/harrier oss v1 270m with the web search query prompt f…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy