Sapiens2 0.4B Pose 308 keypoint top down pose estimation including detailed face (274 keypoints), hand, and foot keypoints. Predictions follow the Sociopticon keypoint format. This repository contains the 0.4B Pose Estimation checkpoint, finetuned from the Sapiens2 0.4B pretrained backbone. Pose is top down — it requires bounding boxes from a person detector. We use RTMDet. 📄 Paper: arXiv:2604.21681 🌐 Project Page: rawalkhirodkar.github.io/sapiens2 💻 Code: github.com/facebookresearch/sapiens2 Model Details Developed by: Meta Model type: Vision Transformer License: Sapiens2 License Task: pose Base model: facebook/sapiens2 pretrain 0.4b Format: safetensors File: sapiens2 0.4b pose.safetensors Quick Start Install the Sapiens2 repo ( pip install e . ), download the checkpoint, and run the demo: See the Pose Estimation guide for details on inputs, outputs, and visualization options. Model Card Field Value Architecture Sapiens2 ViT backbone + Pose Estimation head Backbone parameters 0.398 B Backbone FLOPs 1.260 T Embedding dim 1024 Layers 24 Attention heads 16 Inference resolution 1024 × 768 (H × W) Patch size 16 Sapiens2 Pose Family Model Params FLOPs Embed dim Layers Heads Sapiens2…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy