Model card for PM AI/bi encoder msmarco bert base german Model summary This model can be used for semantic search and documents retrieval to find relevant passages based on a query. It was trained on a machine translated MSMARCO dataset for german with hard negatives and Margin MSE loss . Combining these elements results in a SOTA transformer for asymmetric search. Details are presented below. The model can be easily used with Sentence Transformer library. Training Data The model is based on training with samples from MSMARCO Passage Ranking dataset. It contains about 500.000 questions and 8.8 million passages. The training objective is to identify the relevant passages or answers for an input question. In terms of content, the texts deal with diverse domains. Questions are available as sentences but also keyword based variants can be found. Consequently, models trained on MSMARCO can be used in a variety of domains. The dataset was originally published in English, but has been translated into other languages by researchers with the help of machine translation. To be more specific, "mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset" is used, which contains 13 G…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy