Finetuned SigLIP for Person Visual Descriptions and Reidentification This model is part of the family of SigLIP models finetuned for person visual description and retrieval . Model Details Base model: google/siglip2 base patch16 224 Architecture modifications: none Intended Uses & Limitations Example Applications Person retrieval based on textual or visual descriptions of the person Person image re identification Embedding extraction for retrieval systems Limitations and Bias May inherit biases from the base SigLIP model and training data Not suitable for tasks requiring detailed fine grained recognition without further training Trained on surveillance data; suitable for tasks where a substantial portion of the person is visible Training Loss function Type: Soft contrastive loss with label smoothing; image to image contrastive loss for same identity image pairs. Description: The model is trained to align text and image embeddings using a modified contrastive objective. Instead of relying on hard one hot targets, label smoothing allocates a small probability mass to all other samples in the batch. All embeddings are normalized prior to similarity computation, and the loss is applied…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy