TIPO: Text to Image with text presampling for Prompt Optimization 200M LLaMA arch model trained for TIPO. Tech Report: https://arxiv.org/abs/2411.08127 Introduction In this project, we introduce "TIPO" ( T ext to I mage with text presampling for P rompt O ptimization), an innovative framework designed to significantly enhance the quality and usability of Text to Image (T2I) generative models. TIPO utilizes the Large Language Models (LLMs) to perform "Text Presampling" within the inference pipeline of text to image generative modeling. By refining and extending user input prompts, TIPO enables generative models to produce superior results with minimal user effort, making T2I systems more accessible and effective for a wider range of users. Usage Use updated version of DTG extension (renamed to z tipo extension), current version of z tipo extension support stable diffusion webui, stable diffusion webui forge and ComfyUI. SD Next haven't been tested. https://github.com/KohakuBlueleaf/z tipo extension Model arch and Training This model is LLaMA arch with 200M parameters, the training data is combined version of Danbooru2023, Coyo HD 11M. The total token seen is around 50B tokens. For m…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy