judgebench SigLIP brand judge v1 Part of judgebench: which image QC judges survive optimization pressure? A brand fidelity judge for rhode (beauty brand). A judge = these SigLIP weights + calibration.json (rhode train split centroid + Platt scaling params) in this repo. Base: google/siglip so400m patch14 384 (full SiglipModel , contrastively fine tuned). Training: SupCon contrastive fine tune, rhode positives vs competitor brand negatives (Glossier, ILIA and others), judgebench train split. No violation negatives. Report card (judgebench Phase 1, 2,622 item test set): brand AUC 0.99 , logo masking delta 0.00 (style reader, not a wordmark reader), brand dial Spearman 0.24, near zero violation detection. v1 learned the brand's center , not its boundaries : it is violation blind, and under SRPO gradient pressure it was fully exploited (hacked images scored 0.84). See v2 (+violation negatives) and v3 (hardened) for the ablations. Scoring Score = Platt calibrated cosine similarity between the image embedding and the rhode centroid: Full evaluation protocol, test set construction, and findings: https://github.com/amargupta0428/judgebench.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy