Technique Router (MiniLM) A fine tuned MiniLM classifier that routes image queries to optimal compression techniques for the Headroom SDK. Model Description This model classifies natural language queries about images into one of four optimization techniques: Technique Token Savings Best For transcode ~99% Text extraction, OCR tasks crop 50 90% Region specific queries full low ~87% General understanding preserve 0% Fine details, counting Training Data Base examples : 145 human written queries Expanded dataset : 1,157 examples (via template expansion + synonyms) Split : 85% train, 15% validation Performance Validation Accuracy : 93.7% Model Size : ~128MB Per Class Performance Class Precision Recall F1 Score transcode 0.95 0.92 0.93 crop 0.92 0.97 0.94 preserve 0.97 0.90 0.93 full low 0.89 0.96 0.92 Usage With Headroom SDK Intended Use This model is designed for: Routing image analysis queries to optimal compression techniques Reducing token usage in vision language model applications Enabling cost effective image understanding at scale Limitations English language only Optimized for common image understanding queries May not generalize well to domain specific terminology Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy