๐ง BLIP โ UI Elements Captioning This model is a fine tuned version of Salesforce/blip image captioning base , adapted for captioning UI elements from macOS application screenshots. It is part of the Screen2AX research project focused on improving accessibility using vision based deep learning. ๐ฏ Use Case The model takes an image of a UI icon or element and generates a natural language description (e.g., "Settings icon" , "Play button" , "Search field" ). This helps build assistive technologies such as screen readers by providing textual labels for unlabeled visual components. ๐ Model Architecture Base model: Salesforce/blip image captioning base Architecture: BLIP (Bootstrapping Language Image Pre training) Task: image to text ๐ผ Example ๐ License This model is released under the MIT License . ๐ Related Projects Screen2AX Project Screen2AX HuggingFace Collection โ๏ธ Citation If you use this model in your research, please cite the Screen2AX paper: ๐ MacPaw Research Learn more at https://research.macpaw.com
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy