【ICLR 2024 🔥】LanguageBind: Extending Video Language Pretraining to N modality by Language based Semantic Alignment If you like our project, please give us a star ⭐ on GitHub for latest update. 📰 News [2024.01.27] 👀👀👀 Our MoE LLaVA is released! A sparse model with 3B parameters outperformed the dense model with 7B parameters. [2024.01.16] 🔥🔥🔥 Our LanguageBind has been accepted at ICLR 2024! We earn the score of 6(3)8(6)6(6)6(6) here. [2023.12.15] 💪💪💪 We expand the 💥💥💥 VIDAL dataset and now have 10M video text data . We launch LanguageBind Video 1.5 , checking our model zoo. [2023.12.10] We expand the 💥💥💥 VIDAL dataset and now have 10M depth and 10M thermal data . We are in the process of uploading thermal and depth data on Hugging Face and expect the whole process to last 1 2 months. [2023.11.27] 🔥🔥🔥 We have updated our paper with emergency zero shot results., checking our ✨ results. [2023.11.26] 💥💥💥 We have open sourced all textual sources and corresponding YouTube IDs here. [2023.11.26] 📣📣📣 We have open sourced fully fine tuned Video & Audio , achieving improved performance once again, checking our model zoo. [2023.11.22] We are about to release a fully f…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy