CC Bench: A Cognitive Conflict Benchmark for MLLMs in Safety Critical Visual Inspection CC Bench is a joint medical industrial benchmark for evaluating whether multimodal large language models (MLLMs) remain visually grounded when plausible textual context conflicts with image evidence. The benchmark reorganizes public anomaly datasets into a unified four way multiple choice QA format for high risk visual inspection. This repository currently contains: 4,282 images in total 2,157 normal images and 2,125 anomalous images 11,592 QA instances 3 task families: detection, classification, and localization 6 cognitive conflict types: Expert , History , Lighting , Machine , Noise , and Time The benchmark is designed for research on visual grounding, robustness to misleading context, and reliability in safety critical settings such as medical diagnosis and industrial inspection. Dataset Summary Source datasets and benchmark coverage Source subset Domain Images QA instances Detection Classification Localization : : Br35H Medical 1,200 3,600 Yes Yes Yes LiverCT Medical 911 2,733 Yes Yes Yes DS MVTec Industrial 917 2,751 Yes Yes Yes VisA Industrial 1,254 2,508 Yes No Yes Benchmark protocol For…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy