Depth Anything 3: DA3 SMALL noqa: E501 Model Description DA3 Small model for multi view depth estimation and camera pose estimation. Efficient foundation model with unified depth ray representation. Property Value Model Series Any view Model Parameters 0.08B License Apache 2.0 Capabilities ✅ Relative Depth ✅ Pose Estimation ✅ Pose Conditioning Quick Start Installation Basic Example Command Line Interface Model Details Developed by: ByteDance Seed Team Model Type: Vision Transformer for Visual Geometry Architecture: Plain transformer with unified depth ray representation Training Data: Public academic datasets only Key Insights 💎 A single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization. noqa: E501 ✨ A singular depth ray representation obviates the need for complex multi task learning. Performance 🏆 Depth Anything 3 significantly outperforms: Depth Anything 2 for monocular depth estimation VGGT for multi view depth estimation and pose estimation For detailed benchmarks, please refer to our paper. noqa: E501 Limitations The model is trained on academic datasets and may have limitations on certain domain specific images…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy