Artificial Analysis Long Context Reasoning (AA LCR) Dataset AA LCR includes 100 hard text based questions that require reasoning across multiple real world documents, with each document set averaging ~100k input tokens. Questions are designed such that answers cannot be directly retrieved from documents and must instead be reasoned from multiple information sources. Dataset Development AA LCR was created through a rigorous multi phase process involving several members of the Artificial Analysis research team and more than a dozen undergraduate students who were engaged on a short term contract basis to write and/or validate questions. Document Curation : We selected diverse document sets (company reports, government consultations, legal documents, academic papers) averaging ~100,000 tokens each, representing real materials knowledge workers analyze. Question Creation : Undergraduate students from various disciplines developed questions with access via a dataset development dashboard to non frontier test models to validate question difficulty (GPT 4o mini, Llama 3.1 70B, Gemini 1.5 Flash). These models were specifically chosen to give creators a sense of AI capabilities without acce…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy