Flat Pack Bench Evals 🧪 This repository contains evaluation artifacts for Flat Pack Bench , a CVPR 2026 benchmark for fine grained spatio temporal reasoning in furniture assembly videos. Project page: https://flat pack bench.github.io Benchmark data repo: https://huggingface.co/datasets/justachetan/flat pack bench This repo is intended for analysis and reproducibility. It stores rendered model prompts, generated media artifacts, inference configs, and model responses. The core benchmark questions and media live in the main Flat Pack Bench dataset repo. 📁 Repository Structure Path Contents main result hparam selection/ Main model sweep and prompt/media hyperparameter selection artifacts. qwen 25 vl 72b/ Paper analysis experiments centered on Qwen 2.5 VL 72B. internvl3 78b/ Paper analysis experiments centered on InternVL3 78B. videorefer/ VideoRefer specific questions and outputs, stored separately because the prompt and response format differs from the standard inference format. The standard non VideoRefer experiment layout is: These artifacts were generated by the benchmark inference pipeline in IKEA Manuals at Work/src/benchmark v2/inference.py . At a high level, the pipeline bu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy