Emu3: Next Token Prediction is All You Need Emu3 Team, BAAI Below is the model card of Emu3 Chat model, which is adapted from the original Emu3 model card that you can find here. Model details Model type: Emu3 is an open source multimodal models trained with next token prediction task. By tokenizing images and text into a discrete space, Emu3 is trained as a single transformer from scratch on a mixture of multimodal sequences. It is an auto regressive language model, based on the transformer architecture. Paper or resources for more information: https://github.com/baaivision/Emu3 Highlights Emu3 is capable of generating high quality images following the text input, by simply predicting the next vision token. The model naturally supports flexible resolutions and styles. Emu3 shows strong vision language understanding capabilities to see the physical world and provides coherent text responses. Notably, this capability is achieved without depending on a CLIP and a pretrained LLM. Emu3 simply generates a video causally by predicting the next token in a video sequence, unlike the video diffusion model as in Sora. With a video in context, Emu3 can also naturally extend the video and pred…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy