🌍 GAIA: A global, multimodal, multiscale vision–language dataset for remote sensing image analysis This repository contains the pre trained model weights, associated code, and complete dataset of the paper GAIA: A global, multimodal, multiscale vision–language dataset for remote sensing image analysis. GAIA is a large scale vision language dataset designed to bridge the gap between remote sensing (RS) imagery and natural language understanding. It provides 205,150 image text pairs (41,030 images with 5 synthetic captions each) for advancing RS specific vision language models (VLMs). The dataset spans over 25 years of Earth observations (1998 2024), covering diverse geographic regions, satellite missions, and RS modalities. If you use this work, please cite our paper: A. Zavras, D. Michail, X. X. Zhu, B. Demir and I. Papoutsis, "GAIA: A global, multimodal, multiscale vision–language dataset for remote sensing image analysis," in IEEE Geoscience and Remote Sensing Magazine, doi: 10.1109/MGRS.2025.3650613. Dataset Structure 📂 GAIA has been split into train (70%) , test (20%) , and validation (10%) sets, which are spatio temporally stratified. The dataset splits are provided in JSON…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy