ALOEv2 (multi resolution DINOv3) ALOEv2 is a B cos interpretable DINOv3 model obtained by fine tuning the original ALOE backbones for a further 30k steps. It keeps the original three layer distillation scheme but adds per step multi resolution sampling (224/384/480) and fixes a target layer bug so the even thirds distillation layers include the final transformer block. The original ALOE distilled at a single resolution (224 px), which left features poorly calibrated for the high resolution inputs that dense prediction probes — correspondence, depth, surface normals — actually run on, so feature quality degraded at those resolutions. Training on multiple resolutions removes that train/eval mismatch and restores dense prediction accuracy close to the DINOv3 teacher , while retaining the inherent B cos explanations and holding ImageNet 1k recognition roughly unchanged. This card is shared by the ALOEv2 DINOv3 backbones and their ImageNet 1k linear probe (LP) classifier heads: Kind Repos Backbone rmaser/aloe v2 dinov3 {small,base,large} ImageNet 1k LP rmaser/aloe v2 dinov3 {small,base,large} in1k lp Why ALOEv2 (qualitative difference vs. previous models) Fig 2a — dense correspondence &…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy