Theia The AI Institute Theia is a vision foundation model for robot learning that distills multiple off the shelf vision foundation models trained on varied vision tasks. Theia’s rich visual representations encode diverse visual knowledge, enhancing downstream robot learning. It was introduced in the paper Theia: Distilling Diverse Vision Foundation Models for Robot Learning, which also includes experiments demonstrating that Theia outperforms its teacher models and prior robot learning models using less training data and smaller model sizes. Demo videos can be found on the project page. Model Details The theia tiny patch16 224 cddsv model, uses DeiT Tiny as a backbone, and simulatenously distills CLIP, Depth Anything, DINOv2, Segment Anything and ViT. For more information on usage, please visit the Theia repository. Citation If you use Theia in your research, please use the following BibTeX entry: Usage The pre trained model weights and code released with Theia are available for use under The AI Institute License, reproduced in full below:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy