Trained by Donkey Stereotype Independent Implementation of SPLADE++ Model ( a.k.a splade cocondenser and family ) for the Industry setting. This work stands on the shoulders of 2 robust researches: Naver's From Distillation to Hard Negative Sampling: Making Sparse Neural IR Models More Effective paper and Google's SparseEmbed. Props to both the teams for such a robust work. This is a 2nd iteration in this series. Try V1 here: prithivida/Splade PP en v1 1. What are Sparse Representations and Why learn one? Beginner ? expand this. Expert in Sparse & Dense representations ? feel free skip to next section 2, 1. Lexical search: Lexical search with BOW based sparse vectors are strong baselines, but they famously suffer from vocabulary mismatch problem, as they can only do exact term matching. Here are the pros and cons: ✅ Efficient and Cheap. ✅ No need to fine tune models. ✅️ Interpretable. ✅️ Exact Term Matches. ❌ Vocabulary mismatch (Need to remember exact terms) 2. Semantic Search: Learned Neural / Dense retrievers (DPR, Sentence transformers , BGE models) with approximate nearest neighbors search has shown impressive results. Here are the pros and cons: ✅ Search how humans innately t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy