DSE QWen2 2b MRL V1 DSE QWen2 2b MRL V1 is a bi encoder model designed to encode document screenshots into dense vectors for document retrieval. The Document Screenshot Embedding (DSE) approach captures documents in their original visual format, preserving all information such as text, images, and layout, thus avoiding tedious parsing and potential information loss. DSE aims to provide a generalizable embedding model for Text, PDF documents, Webpage, Slides retrieval. For example, DSE QWen2 2b MRL V1 achieves 85.8 nDCG@5 on ViDoRE leaderboard. Note: QWen vision encoder may take high GPU memory if the input image is large. Adjust 'resized height':680 , 'resized width':680 (see below) to fit VRAM based on GPU resources. How to Use the Model To support better effectiveness efficiency trade off, this checkpoint is trained to support: 1. Flexible representation dimension. 2. Flexible input image size. Load the Model and Processor Encode Text Query Encode Document Screenshot Compute Similarity Encode Document Text This DSE checkpoint is warm up with Tevatron/msmarco passage aug , thus the model can also effectively encode document as text input. Citation If you find this checkpoint is he…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy