RealText V2: A Large Scale Multilingual Document Forgery Analysis Benchmark 💾 Dataset Description RealText V2 is a large scale multilingual document benchmark dataset purpose built for multilingual text image forgery analysis, pioneering in both scale and annotation depth. Key Features 20K+ images : A large scale benchmark, surpassing existing document forgery analysis datasets by orders of magnitude 6 languages : English, Chinese, Arabic, Thai, Malay, and Indonesian — spanning Latin, logographic, Arabic, and Thai script systems, each presenting unique forgery analysis challenges 6 domains : Finance, education, healthcare, live streaming, e commerce, and natural scenes Multi granularity forgery : Character level, word level, and semantic level tampering Multi source samples : Real world and AIGC synthesized forgery samples covering diverse generation pipelines Rich multi task annotations : Pixel level localization masks, tampering type labels, and expert level natural language explanations Competition Timeline ACM MM 2026 MGC: GenText Forensics: Challenge on Explainable Forensics and Adversarial Generation for Text Centric Images https://www.codabench.org/competitions/15805/ Phase…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy