BridgeTower large itm mlm itc model The BridgeTower model was proposed in "BridgeTower: Building Bridges Between Encoders in Vision Language Representative Learning" by Xiao Xu, Chenfei Wu, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan. The model was pretrained on English language using masked language modeling (MLM) and image text matching (ITM)objectives. It was introduced in this paper and first released in this repository. BridgeTower got accepted to AAAI'23. Model description The abstract from the paper is the following: Vision Language (VL) models with the Two Tower architecture have dominated visual language representation learning in recent years. Current VL models either use lightweight uni modal encoders and learn to extract, align and fuse both modalities simultaneously in a deep cross modal encoder, or feed the last layer uni modal representations from the deep pre trained uni modal encoders into the top cross modal encoder. Both approaches potentially restrict vision language representation learning and limit model performance. In this paper, we propose BridgeTower, which introduces multiple bridge layers that build a connection between the top layers of uni mo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy