Substream Recollection A controlled benchmark for substream membership recall in long context VLMs and LLMs. Each row is a (stream, probe, label) tuple: the model sees a long input stream and a short probe, and must answer "yes" or "no" — did the probe occur inside the stream? The dataset is organized into four top level configs keyed by modality + source: config rows content text 7,640 text modality questions for the synthetic substream benchmark. synthetic video 6,065 rendered synthetic substream videos. easyhuman 672 rendered 3 belt EasyHuman videos (224 video rows at L=256) plus their text modality counterparts (448 text rows at L=256 + L=1024). The modality column distinguishes text vs video . Pattern based, so no entropy ground truth on this config. natural video 1,028 EPIC Kitchens 100 derived clips and SoccerNet provenance metadata. Directory layout Inside each synthetic video/videos/L frames/ / you'll find the parent videos (e.g. video 1 v0.mp4 ) plus a clips/ subfolder with the probe clips. EasyHuman uses a flatter layout with no inner bucket directory. Loading Each config ships in three formats so downstream consumers can pick whichever is most convenient: /questions.par…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy