Sapiens2 5B Pose 308 keypoint top down pose estimation including detailed face (274 keypoints), hand, and foot keypoints. Predictions follow the Sociopticon keypoint format. This repository contains the 5B Pose Estimation checkpoint, finetuned from the Sapiens2 5B pretrained backbone. Pose is top down — it requires bounding boxes from a person detector. We use RTMDet. 📄 Paper: arXiv:2604.21681 🌐 Project Page: rawalkhirodkar.github.io/sapiens2 💻 Code: github.com/facebookresearch/sapiens2 Model Details Developed by: Meta Model type: Vision Transformer License: Sapiens2 License Task: pose Base model: facebook/sapiens2 pretrain 5b Format: safetensors File: sapiens2 5b pose.safetensors Quick Start Install the Sapiens2 repo ( pip install e . ), download the checkpoint, and run the demo: See the Pose Estimation guide for details on inputs, outputs, and visualization options. Model Card Field Value Architecture Sapiens2 ViT backbone + Pose Estimation head Backbone parameters 5.071 B Backbone FLOPs 15.722 T Embedding dim 2432 Layers 56 Attention heads 32 Inference resolution 1024 × 768 (H × W) Patch size 16 Sapiens2 Pose Family Model Params FLOPs Embed dim Layers Heads Sapiens2 0.4B 0.39…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy