TACIT Benchmark v0.1.0 Transformation Aware Capturing of Implicit Thought A programmatic visual reasoning benchmark for evaluating generative and discriminative capabilities of multimodal models across 10 tasks and 6 reasoning domains. Author: Daniel Nobrega Medeiros arXiv paper GitHub Overview TACIT presents visual puzzles that require genuine spatial, logical, and structural reasoning — not pattern matching on text. Each puzzle is generated programmatically with deterministic seeding, ensuring full reproducibility. Evaluation is programmatic (no LLM as judge): solutions are verified through computer vision algorithms (pixel sampling, SSIM, BFS path detection, color counting). Key Features 6,000 puzzles across 10 tasks and 3 difficulty levels Dual track evaluation : generative (produce a solution image) and discriminative (select from candidates) Multi resolution : every puzzle rendered at 512px, 1024px, and 2048px Deterministic : seeded generation (seed=42) for exact reproducibility Programmatic verification : CV based solution checking, no subjective evaluation Task Examples All examples below show medium difficulty puzzles at 512px resolution. 01 — M…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy