Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation Yichen Zhang 1 , Da Peng 2 , Zonghao Guo 1† , Zijian Zhang 3 , Xuesong Yang 3 , Tong Sun 3 , Shichu Sun 3 , Yidan Zhang 3 , Yanghao Li 1 , Haiyan Zhao 1 , Wang Xu 1 , Qi Shi 1 , Yangang Sun 1 , Chi Chen 1 , Shuo Wang 1 , Yukun Yan 1 , Xu Han 1 , Qiang Ma 1 , Wei Ke 2 , Liang Wang 3 , Zhiyuan Liu 1 , Maosong Sun 1 1 Tsinghua University, 2 Xi'an Jiaotong University, 3 University of Chinese Academy of Sciences \ Equal contribution † Corresponding author 🌟 What is Cheers ? A recent cutting edge topic in multimodal modeling is to unify visual comprehension and generation within a single model. However, the two tasks demand mismatched decoding regimes and visual representations, making it non trivial to jointly optimize within a shared feature space. In this work, we present Cheers , a unified multimodal model that decouples patch level details from semantic representations, thereby stabilizing semantics for multimodal understanding and improving fidelity for image generation via gated detail residuals. Cheers includes three key components: (i) a unified vision tokenize…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy