First version of JaColBERTv2. Weights might be updated in the next few days. Current early checkpoint is fully functional and outperforms multilingual e5 large, BGE M3 and JaColBERT in early results, but full evaluation TBD. Intro There is currently no JaColBERTv2 technical report. For an overall idea, you can refer to the JaColBERTv1 arXiv Report If you just want to check out how to use the model, please check out the Usage section below! Welcome to JaColBERT version 2, the second release of JaColBERT, a Japanese only document retrieval model based on ColBERT. JaColBERTv2 is a model that offers very strong out of domain generalisation. Having been only trained on a single dataset (MMarco), it reaches state of the art performance. JaColBERTv2 was initialised off JaColBERTv1 and trained using knowledge distillation with 31 negative examples per positive example. It was trained for 250k steps using a batch size of 32. The information on this model card is minimal and intends to give a quick overview! It'll be updated once benchmarking is complete and a longer report is available. Why use a ColBERT like approach for your RAG application? Most retrieval methods have strong tradeoffs: T…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy