📚 Collection 📝 Blog LateOn State of the Art ColBERT Retrieval Model by LightOn DenseOn LateOn PyLate FastPLAID About the LateOn / DenseOn Family State of the art retrieval is increasingly dominated by closed models, either hidden behind APIs or trained on undisclosed data. This blocks reproducibility, prevents study of possible data leakage, and gatekeeps progress to a handful of private labs. We thus decided to gather and curate a large amount of data and explore various mixtures. We release all the data used in our explorations: Gathered pre training data, 1.4B query documents pairs alongside annotations used for non destructive filtering (structural filtering, deduplication, cross encoder pair relevancy) Best pre training mixture found with already applied filters Fine tuning datasets with query, positive and 2048 mined documents alongside their scores for 1.88M samples. Based on our findings, we trained LateOn (multi vector/ColBERT) and DenseOn (single vector/dense) models on a proprietary Apache 2.0 compatible training dataset and release those models as well. Both are built on the ModernBERT backbone at 149M parameters, a size we believe sits at the sweet spot: large enough…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy