IndustryBench MIPU: Benchmarking Multi Image Attribute Value Extraction for Industrial Products Multi Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real world industrial product images. Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench MIPU tests whether MLLMs can reliably recover them through four challenges: text recognition, visual reasoning, domain knowledge, and cross image evidence integration. 🚧 Work in Progress — We are still actively updating the benchmark to further improve its quality. A new version is expected to be released by late June 2026 . Main Results Model Multi P Multi R Multi F1 Gemini 3.1 Pro 93.8 49.9 65.1 Qwen 3.5 397B A17B 88.2 48.6 62.7 GPT 5.4 86.3 46.6 60.5 Qwen 3.5 Plus 88.1 45.4 59.9 Claude Opus 4.6 88.2 42.3 57.2 Key finding: Models achieve high precision (86–94%) but the best recovers only 49.9% of product level attributes. Multi image completeness is the core bottleneck. Dataset Summary IndustryBench MIPU is the first large scale benchmark for multi image industrial product understand…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy