FoodLensVN — Vietnamese Food VQA A small but clean Vietnamese language Visual Question Answering dataset over 20 canonical Vietnamese dishes . Built as the Phase 1 corpus for the FoodLensVN project (final year Deep Learning report). Contents 5,572 (image, question, answer) rows — 4,460 train / 632 val / 480 test . 299 unique source images + 897 augmented variants (3× per source) → 1,196 image refs per variant . Two image variants: raw/ (original aspect ratio, max edge ≤ 1024) and squared/ (center padded to a square — recommended for training). Splits are image disjoint at the image id level, stratified by dish . Answers are full Vietnamese sentences (≈7–9 words, ≤10 after canonicalization). Layout Row schema Locked dish set (20) pho , bun bo hue , banh mi , com tam , bun cha , goi cuon , cha gio , banh xeo , mi quang , hu tieu , banh cuon , bun thit nuong , cao lau , bot chien , banh khot , xoi xeo , chao long , bun dau mam tom , bun mam , banh canh . Question types (6) yes no , counting , recognition , attribute , spatial , reasoning — roughly balanced per split. See ANNOTATOR GUIDE.md in the repo for details. Loading On Kaggle, set the HF TOKEN notebook secret (read scoped is eno…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy