ViT Base Violence Detection Model Description This is a Vision Transformer (ViT) model fine tuned for violence detection. The model is based on google/vit base patch16 224 in21k and has been trained on the Real Life Violence Situations dataset from Kaggle to classify images into violent or non violent categories. Intended Use The model is intended for use in applications where detecting violent content in images is necessary. This can include: Content moderation Surveillance Parental control software Model accuracy Test accuracy for Vit Base = 98.80% Loss = 0.20038144290447235 How to Use Here is an example of how to use this model for image classification: python import torch from transformers import ViTForImageClassification, ViTFeatureExtractor from PIL import Image Load the model and feature extractor model = ViTForImageClassification.from pretrained('jaranohaal/vit base violence detection') feature extractor = ViTFeatureExtractor.from pretrained('jaranohaal/vit base violence detection') Load an image image = Image.open('image.jpg') Preprocess the image inputs = feature extractor(images=image, return tensors="pt") Perform inference with torch.no grad(): outputs = model( inputs)…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy