🚧 Update [x] (Sep 29th, 2025) We updated our paper, where we removed some in efficient and high cost samples. We also added a sub sample of DetectiveQA. [x] (July 7th, 2025) We released the initial version of our datasets. [x] (July 22nd, 2025) We modify the datasets slightly, adding the keypoints in LRU and change the into . The is only used in Longmemeval task. [x] (July 26th, 2025) We fixed bug on . [x] (Aug.5th, 2025) We removed the and some other datasets not used in main experiments. We will release a subset for ablation study in future. ⚙️ MemoryAgentBench: Evaluating Memory in LLM Agents via Incremental Multi Turn Interactions This repository contains the MemoryAgentBench dataset, designed for evaluating the memory capabilities of LLM agents. 📄 Paper: https://arxiv.org/pdf/2507.05257 💻 Code: https://github.com/HUST AI HYZ/MemoryAgentBench MemoryAgentBench is a unified benchmark framework for comprehensively evaluating the memory capabilities of LLM agents: through four core competencies (Accurate Retrieval, Test Time Learning, Long Range Understanding, and Conflict Resolution) and incremental multi turn interaction design, it reveals existing limitations and shortcomings…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy