pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs pdfQA is a structured benchmark collection for document level question answering and PDF understanding research. The dataset is organized to support: Raw document processing research Structured extraction pipelines Retrieval augmented QA End to end document reasoning systems It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies. Dataset Structure ⚠️ The QA annotations are released as a separate Hugging Face dataset: https://huggingface.co/datasets/pdfqa/pdfQA Annotations (Dataset ID: pdfqa/pdfQA Annotations ) This repository follows a strict hierarchical layout: Categories real pdfQA/ — Real world benchmark datasets syn pdfQA/ — Synthetic benchmark datasets Types Each dataset contains three file type folders: 01.1 Input Files Non PDF/ — Original source formats (e.g., xlsx, epub, htm, tex, txt) 01.2 Input Files PDF/ — Original PDF files 01.3 Input Files CSV/ — Structured tabular representations Datasets Each type folder contains subfolders for individual datasets. Supported datasets include: Real world Datasets ClimateFinanceBench/ Clim…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy