ALOE: Align Once to Explain Using DINOv3? A newer ALOEv2 multi resolution DINOv3 model is available and recommended — see rmaser/aloe v2 dinov3 {small,base,large} (and the matching in1k lp classifier heads). ALOEv2 substantially improves dense correspondence and depth while keeping inherent B cos explanations. This repository contains one of the published ALOE vision backbones or ImageNet 1k linear probe classifiers from the accepted CVPR 2026 poster "Align Once to Explain: Feature Alignment for Scalable B cosification of Foundational Vision Transformers" . ALOE converts a frozen ViT style foundation model into an inherently interpretable B cos counterpart through a one time, label free feature alignment stage. The aligned model is meant to be used as a drop in visual backbone: it keeps strong downstream representations while exposing model inherent B cos explanations from the network itself. ALOE stays within a fraction of a point of the original foundation models on ImageNet 1k linear probe accuracy while lifting Grid PG localization far above the teachers' best post hoc explainers. What ALOE Does Starts from a frozen teacher encoder such as supervised ViT B/16, DINOv3, or SigLIP…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy