judgebench SigLIP brand judge v3 (hardened) Part of judgebench: which image QC judges survive optimization pressure? A brand fidelity judge for rhode (beauty brand). A judge = these SigLIP weights + calibration.json (rhode train split centroid + Platt scaling params) in this repo. Base: google/siglip so400m patch14 384 (full SiglipModel , contrastively fine tuned). Training: the v1 recipe plus SRPO hacked images (outputs of the gradient attack on v1) folded in as a third negative class. Report card (judgebench Phase 1, 2,622 item test set): real rhode brand AUC 0.997 (v1 0.99); real rhode scores 0.97 vs competitors 0.06. The arms race result: the seen attack is fully defeated (SRPO hack images: v1 scored 0.84 → v3 scores 0.00) while the unseen DPO attack is only dampened (0.47 → 0.28). Hardening against an observed exploit works; novel attacks still get partial traction. Scoring Score = Platt calibrated cosine similarity between the image embedding and the rhode centroid: Full evaluation protocol, test set construction, and findings: https://github.com/amargupta0428/judgebench.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy