PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre process the datasets into a uniform format and write several task specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks: Image to Text (I2T) retrieval These tasks evaluate the ability of a retriever to find relevant documents associated with an input image. Component tasks are WIT, IGLUE en, KVQA, and CC3M. Question to Text (Q2T) retrieval This task is based on MSMARCO and is included to assess whether multi modal retrievers retain their ability in text only retrieval after any retraining for images. Image & Question to Text (IQ2T) retrieval This is the most challenging task which requires joint understanding of questions and images for accurate retrieval. It consists of these subtasks: OVEN, LLaVA, OKVQA, Infoseek and E VQA. Paper or resources for more information: Paper: https://arxiv.org/abs/2402.08327 Project Page: https://preflmr.github.io…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy