LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long context Multitasks π Project Page: https://longbench2.github.io π» Github Repo: https://github.com/THUDM/LongBench π Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long context problems requiring deep understanding and reasoning across real world multitasks. LongBench v2 has the following features: (1) Length : Context length ranging from 8k to 2M words, with the majority under 128k. (2) Difficulty : Challenging enough that even human experts, using search tools within the document, cannot answer correctly in a short time. (3) Coverage : Cover various realistic scenarios. (4) Reliability : All in a multiple choice question format for reliable evaluation. To elaborate, LongBench v2 consists of 503 challenging multiple choice questions, with contexts ranging from 8k to 2M words, across six major task categories: single document QA, multi document QA, long in context learning, long dialogue history understanding, code repo understanding, and long structured data understanding. To ensure the breadth and the practicality, we collect data from nearlβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy