Banana Merged A synthetic multi page visual question answering dataset with hard negatives, designed for fine tuning visual document retrievers like ColFlor and ColPali. Dataset Summary Banana Merged contains 1,100 training samples and 10,054 images (positive pages + hard negative variants). Each sample pairs a multi page analytical query with a set of document images that collectively contain the answer, plus one or more hard negative documents that look visually and lexically similar but would be the wrong retrieval result. The dataset was generated by the Nano Banana Pro pipeline using Gemini 3 Pro Image Preview for image generation and editing. The pipeline takes a seed document image (sourced from llamaindex/vdr multilingual train and similar corpora), synthesizes a multi page document around it, generates a query that requires reading across all pages to answer, then produces hard negative variants by deliberately modifying specific facts or attributes while preserving the visual style. Hard negatives are critical for training retrieval models to distinguish between documents that share surface level keywords and layout but differ in the specific information needed to answer…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy