Model Details: DPT Large (also known as MiDaS 3.0) Dense Prediction Transformer (DPT) model trained on 1.4 million images for monocular depth estimation. It was introduced in the paper Vision Transformers for Dense Prediction by Ranftl et al. (2021) and first released in this repository. DPT uses the Vision Transformer (ViT) as backbone and adds a neck + head on top for monocular depth estimation. The model card has been written in combination by the Hugging Face team and Intel. Model Detail Description Model Authors Company Intel Date March 22, 2022 Version 1 Type Computer Vision Monocular Depth Estimation Paper or Other Resources Vision Transformers for Dense Prediction and GitHub Repo License Apache 2.0 Questions or Comments Community Tab and Intel Developers Discord Intended Use Description Primary intended uses You can use the raw model for zero shot monocular depth estimation. See the model hub to look for fine tuned versions on a task that interests you. Primary intended users Anyone doing monocular depth estimation Out of scope uses This model in most cases will need to be fine tuned for your particular task. The model should not be used to intentionally create hostile or a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy