🗺️ UAVReason Depth Depth maps and depth statistics for UAV native multimodal reasoning and generation 📌 News Paper: Can Vision Language Models Think from the Sky? Unifying UAV Reasoning and Generation arXiv: arXiv:2604.05377 VQA / caption / generation annotations: jarvissun/UAVReason vqa This dataset is released as part of UAVReason , introduced in the paper above. Please cite the paper if you use this dataset. 🧭 Overview UAVReason Depth provides depth maps and depth related metadata for the UAVReason benchmark. It is designed to support geometry aware aerial perception, depth aware visual question answering, and cross modal generation across RGB, depth, semantic segmentation, and language. Depth information is important for UAV view reasoning because aerial images often contain small objects, compressed perspective, weak texture cues, and ambiguous spatial layouts. The depth maps in this repository provide dense geometric priors for studying whether multimodal models can ground their reasoning in the 3D structure of aerial scenes. The full UAVReason benchmark aligns RGB imagery , depth maps , semantic segmentation masks , captions , and question answer pairs in a consistent aer…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy