Overview This model implements a novel approach to multi reference video generation using Multiple Subject Reference (MSR) . Instead of introducing additional encoder branches or fusion modules, we transform multiple static reference images into a pseudo video sequence that shares the same representation space as the target video. Usage This LoRA requires the ComfyUI Licon MSR plugin for ComfyUI. A sample workflow is included in the model files for easy testing and experimentation. Key Features Multi Reference Visual Memory Token level reference preservation : Multiple reference images are encoded as video latents, preserving fine grained visual information at token level rather than compressing into a single embedding Native self attention retrieval : The target video tokens directly access reference tokens through the model's existing self attention mechanism—no new architectural components needed In context conditioning : References serve as "visual memory" within the main token sequence, not as external conditioning inputs Flexible Reference Composition 2 to 5 reference images : Supports varying numbers of reference inputs with increasing complexity Complementary semantic roles…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy