MMAD: The First Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection 💡 This dataset is the full version of MMAD Content :Containing both questions, images, and captions. Questions : All questions are presented in a multiple choice format with manual verification, including options and answers. Images :Images are collected from the following links: DS MVTec , MVTec AD , MVTec LOCO , VisA , GoodsAD. We retained the mask format of the ground truth to facilitate future evaluations of the segmentation performance of multimodal large language models. Captions :Most images have a corresponding text file with the same name in the same folder, which contains the associated caption. Since this is not the primary focus of this benchmark, we did not perform manual verification. Although most captions are of good quality, please use them with caution. 👀 Overview In the field of industrial inspection, Multimodal Large Language Models (MLLMs) have a high potential to renew the paradigms in practical applications due to their robust language capabilities and generalization abilities. However, despite their impressive problem solving skills in many dom…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy