Depth Anything 3: DA3MONO LARGE noqa: E501 Model Description DA3 Monocular Large model for high quality relative monocular depth estimation. Unlike disparity based models (e.g., Depth Anything 2), it directly predicts depth, resulting in superior geometric accuracy. Property Value Model Series Monocular Depth Parameters 0.35B License Apache 2.0 Capabilities ✅ Relative Depth ✅ Sky Segmentation Quick Start Installation Basic Example Command Line Interface Model Details Developed by: ByteDance Seed Team Model Type: Vision Transformer for Visual Geometry Architecture: Plain transformer with unified depth ray representation Training Data: Public academic datasets only Key Insights 💎 A single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization. noqa: E501 ✨ A singular depth ray representation obviates the need for complex multi task learning. Performance 🏆 Depth Anything 3 significantly outperforms: Depth Anything 2 for monocular depth estimation VGGT for multi view depth estimation and pose estimation For detailed benchmarks, please refer to our paper. noqa: E501 Limitations The model is trained on academic datasets and may h…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy