Depth Anything 3: DA3 GIANT noqa: E501 Model Description DA3 Giant model for multi view depth estimation, camera pose estimation, and 3D Gaussian estimation. This is the flagship foundation model with unified depth ray representation. Property Value Model Series Any view Model Parameters 1.15B License CC BY NC 4.0 ⚠️ Non commercial use only due to CC BY NC 4.0 license. Capabilities ✅ Relative Depth ✅ Pose Estimation ✅ Pose Conditioning ✅ 3D Gaussians Quick Start Installation Basic Example Command Line Interface Model Details Developed by: ByteDance Seed Team Model Type: Vision Transformer for Visual Geometry Architecture: Plain transformer with unified depth ray representation Training Data: Public academic datasets only Key Insights 💎 A single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization. noqa: E501 ✨ A singular depth ray representation obviates the need for complex multi task learning. Performance 🏆 Depth Anything 3 significantly outperforms: Depth Anything 2 for monocular depth estimation VGGT for multi view depth estimation and pose estimation For detailed benchmarks, please refer to our paper. noqa: E501 Limi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy