TerraMind 1.0 large TerraMind is the first multimodal any to any generative foundation model for Earth Observation jointly developed by IBM, ESA, and Forschungszentrum Jülich. Architecture TerraMind uses a dual scale transformer based encoder decoder architecture, simultaneously processing pixel level and token level data. The model was pre trained on 500B tokens from 9M spatiotemporally aligned multimodal samples from the TerraMesh dataset. Modality specific patch embeddings allow direct processing of raw inputs, while modality specific FSQ VAEs are used for image tokenization. For sequence like modalities such as coordinates, an adapted WordPiece tokenizer is employed. During pre training, TerraMind leverages masked token reconstruction, learning complex cross modal correlations to generate high quality latent representations. Evaluation We benchmarked TerraMind against other geospatial foundation models using the PANGAEA benchmark. TerraMind consistently achieved state of the art performance, surpassing existing models in various downstream tasks such as land use segmentation, water body mapping, and vegetation assessments. The evaluation highlights its effectiveness in handling…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy