Depth Anything V2 Base – Transformers Version Depth Anything V2 is trained from 595K synthetic labeled images and 62M+ real unlabeled images, providing the most capable monocular depth estimation (MDE) model with the following features: more fine grained details than Depth Anything V1 more robust than Depth Anything V1 and SD based models (e.g., Marigold, Geowizard) more efficient (10x faster) and more lightweight than SD based models impressive fine tuned performance with our pre trained models This model checkpoint is compatible with the transformers library. Depth Anything V2 was introduced in the paper of the same name by Lihe Yang et al. It uses the same architecture as the original Depth Anything release, but uses synthetic data and a larger capacity teacher model to achieve much finer and robust depth predictions. The original Depth Anything model was introduced in the paper Depth Anything: Unleashing the Power of Large Scale Unlabeled Data by Lihe Yang et al., and was first released in this repository. Online demo. Model description Depth Anything V2 leverages the DPT architecture with a DINOv2 backbone. The model is trained on ~600K synthetic labeled images and ~62 million…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy