Depth Anything 3: DA3 LARGE noqa: E501 Model Description DA3 Large model for multi view depth estimation and camera pose estimation. Foundation model with unified depth ray representation. Property Value Model Series Any view Model Parameters 0.35B License CC BY NC 4.0 Capabilities ✅ Relative Depth ✅ Pose Estimation ✅ Pose Conditioning Quick Start Installation Basic Example Command Line Interface Model Details Developed by: ByteDance Seed Team Model Type: Vision Transformer for Visual Geometry Architecture: Plain transformer with unified depth ray representation Training Data: Public academic datasets only Key Insights 💎 A single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization. noqa: E501 ✨ A singular depth ray representation obviates the need for complex multi task learning. Performance 🏆 Depth Anything 3 significantly outperforms: Depth Anything 2 for monocular depth estimation VGGT for multi view depth estimation and pose estimation For detailed benchmarks, please refer to our paper. noqa: E501 Limitations The model is trained on academic datasets and may have limitations on certain domain specific images noqa: E5…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy