polish reranker roberta v3 This model is designed to serve as a Polish reranker in retrieval augmented generation (RAG) pipelines. It is a general purpose reranker that delivers strong performance across various document types and domains. It is the successor to sdadas/polish reranker roberta v2 and has the following key features: The model is based on a new version of polish roberta that supports long contexts of up to 8192 tokens . It was trained using knowledge distillation, with a LLM based reranker (BAAI/bge reranker v2.5 gemma2 lightweight) as the teacher model. The training involved the same RankNet loss function and largely the same training corpus as the previous version. The only major change is the addition of a new dataset comprising approximately 40,000 passages from the public administration domain, sourced from Polish government and local municipality websites. Questions for these passages were generated synthetically using DeepSeek V3 0324 . The model achieves significantly better results on long context reranking tasks than the previous version. On tasks involving short texts, it maintains performance similar to sdadas/polish reranker roberta v2 . This is an effici…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy