Dataset Description Dataset Summary The NeedleBench dataset is a part of the OpenCompass project, designed to evaluate the capabilities of large language models (LLMs) in processing and understanding long documents. It includes a series of test scenarios that assess models' abilities in long text information extraction and reasoning. The dataset is structured to support tasks such as single needle retrieval, multi needle retrieval, multi needle reasoning, and ancestral trace challenges. Supported Tasks and Primary Languages Single Needle Retrieval Task (S RT) : Extracting a single key piece of information from a long text. Multi Needle Retrieval Task (M RT) : Retrieving multiple related pieces of information from long texts. Multi Needle Reasoning Task (M RS) : Extracting and utilizing multiple key pieces of information for comprehensive understanding. Ancestral Trace Challenge (ATC) : Handling multi layer logical challenges in real long texts. The dataset supports multiple languages, including English and Chinese, as indicated by the presence of files like multi needle reasoning en.json and multi needle reasoning zh.json . Potential Use Cases The NeedleBench dataset can be used to…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy