LLMDet (large variant) LLMDet model was proposed in LLMDet: Learning Strong Open Vocabulary Object Detectors under the Supervision of Large Language Models by Shenghao Fu, Qize Yang, Qijie Mo, Junkai Yan, Xihan Wei, Jingke Meng, Xiaohua Xie, Wei Shi Zheng. LLMDet improves upon the MM Grounding DINO and Grounding DINO by co training the model with a large language model. You can find all the LLMDet checkpoints under the LLMDet collection. Note that these checkpoints are inference only they do not include LLM which was used for training. The inference is identical to that of MM Grounding DINO. Intended uses You can use the raw model for zero shot object detection. Here's how to use the model for zero shot object detection: Training Data This model was trained on: Objects365v1 Open Images v6 GOLD G GroundingCap 1M Evaluation results Here's a table of LLMDet models and their performance on LVIS (results from official repo): Model Pre Train Data MiniVal APr MiniVal APc MiniVal APf MiniVal AP Val1.0 APr Val1.0 APc Val1.0 APf Val1.0 AP llmdet tiny (O365,GoldG,GRIT,V3Det) + GroundingCap 1M 44.7 37.3 39.5 50.7 34.9 26.0 30.1 44.3 llmdet base (O365,GoldG,V3Det) + GroundingCap 1M 48.3 40.8 43…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy