BLINK: Multimodal Large Language Models Can See but Not Perceive π Homepage π» Code π Paper π arXiv π Eval AI This page contains the benchmark dataset for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive" Introduction We introduce BLINK, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the BLINK tasks can be solved by humans βwithinβ¦ See the full description on the dataset page: https://huggingface.co/datasets/BLINK Benchmark/BLINK.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy