🔥 Mantis (TMLR 2024) Paper Website Github Models Demo Wandb Summary Mantis is an LLaMA 3 based LMM with interleaved text and image as inputs , train on Mantis Instruct under academic level resources (i.e. 36 hours on 16xA100 40G). Mantis is trained to have multi image skills including co reference, reasoning, comparing, temporal understanding. Mantis reaches the state of the art performance on five multi image benchmarks (NLVR2, Q Bench, BLINK, MVBench, Mantis Eval), and also maintain a strong single image performance on par with CogVLM and Emu2. Multi Image Performance Models Size Format NLVR2 Q Bench Mantis Eval BLINK MVBench Avg : : : : : : : : : : : : : : : : GPT 4V sequence 88.80 76.52 62.67 51.14 43.50 64.5 Open Source Models Random 48.93 40.20 23.04 38.09 27.30 35.5 Kosmos2 1.6B merge 49.00 35.10 30.41 37.50 21.62 34.7 LLaVA v1.5 7B merge 53.88 49.32 31.34 37.13 36.00 41.5 LLava V1.6 7B merge 58.88 54.80 45.62 39.55 40.90 48.0 Qwen VL Chat 7B merge 58.72 45.90 39.17 31.17 42.15 43.4 Fuyu 8B merge 51.10 49.15 27.19 36.59 30.20 38.8 BLIP 2 13B merge 59.42 51.20 49.77 39.45 31.40 46.2 InstructBLIP 13B merge 60.26 44.30 45.62 42.24 32.50 45.0 CogVLM 17B merge 58.58 53.20 45.16…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy