π $\\pi^3$: Scalable Permutation Equivariant Visual Geometry Learning $\\pi^3$ reconstructs visual geometry without a fixed reference view, achieving robust, state of the art performance. β¨ Overview We introduce $\\pi^3$ (Pi Cubed), a novel feed forward neural network that revolutionizes visual geometry reconstruction by eliminating the need for a fixed reference view . Traditional methods, which rely on a designated reference frame, are often prone to instability and failure if the reference is suboptimal. In contrast, $\\pi^3$ employs a fully permutation equivariant architecture. This allows it to directly predict affine invariant camera poses and scale invariant local point maps from an unordered set of images, breaking free from the constraints of a reference frame. This design makes our model inherently robust to input ordering and highly scalable . A key emergent property of our simple, bias free design is the learning of a dense and structured latent representation of the camera pose manifold. Without complex priors or training schemes, $\\pi^3$ achieves state of the art performance π on a wide range of tasks, including camera pose estimation, monocular/video depth estimatβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy