JoyAI Echo ๐ฌ Pushing the Frontier of Long Video Generation Official model weights for minute level multi shot audio video generation with a distilled DMD generator, paired cross modal memory, and story level consistency. For academic research and non commercial use only. ๐ Paper ๐ Project Page ๐ป Inference Code ๐งฌ Model ๐ Usage ๐ Results ๐ Citation Model Summary JoyAI Echo is a long form, multi shot, audio video generation framework that breaks the barriers of error accumulation, weak temporal coherence, and prohibitive latency in long video generation. A cross modal audio visual memory bank preserves character appearance and voice timbre consistently over five minute videos, while a post training pipeline combining memory based reinforcement learning with distribution matching distillation (DMD) delivers a 7.5ร inference speedup without sacrificing quality. JoyAI Echo decisively outperforms HappyOyster (directing mode) on long form generation and even surpasses the short video specialist Wan 2.6 on human centric tasks. This repository hosts the released checkpoint . Inference code is released separately โ see the Usage section. Model Details Developed by: Echo Team @ Joy Futโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy