🚀 Singularity-LTX-2.3_OmniCine_V1 Official Release
This is not just a standard fine-tune; it is a fundamental restructuring of the LTX-Video (2.3) generation logic.
I am thrilled to present the official release of LTX2.3 Singularity to the community. This comprehensive optimization framework focuses heavily on Image-to-Video (I2V), First & Last Frame Control, and Reference-to-Video generation. Although it has currently undergone only nearly 100,000 steps (calculated by gradient accumulation), its enhancements in physical consistency, dynamic motion, and cinematic expression have already far exceeded expectations.
🌟 Key Improvements
- 🦴 Limbs & Anatomy Evolution: Specifically optimized to fix the common degradation of fingers and toes, drastically reducing anatomy warping and artifacts during fast movements.
- 🎬 Injecting Shot Continuity: Achieved precise timeline-based shot and camera cuts controlled directly via text prompts (0-5s logical segments), saying goodbye to erratic, randomized framing.
- 🗣️ Elimination of "AI Stiffness": Significantly enhanced facial expressiveness during speech, deeply optimized lip-syncing, and natively eliminated the rigid, burned-in subtitles frequently generated by the base model.
- ⚖️ Physical Consistency: Improved the structural integrity of characters and environments during high-speed actions, suppressing chaotic "twisting/morphing" and aligning motions with real-world physics.
- 🎨 Flawless Anime Compatibility: Integrated a high-quality Anime training dataset, allowing the model to seamlessly adapt across diverse styles including 2D anime, 3D CGI, and hyper-realism.
- 🌪️ Extreme Dynamic Range: Delivers stellar performance in high-action sequences like running and combat sports. Simultaneously, visual effects for cyberpunk themes, transformations, magic casting, and monster rendering have been massively amplified.
- 🖼️ Revolutionary Reference Image Control: Upgraded the "Reference-to-Video" capability. No longer bound to rigid first-frame constraints, the model intelligently extracts character features and artistic styles from the reference image, generating entirely new angles and compositions based on your prompts.
Effect Demonstration
📊 Performance Showreel
| Evaluation Dimension | Performance Characteristics |
|---|---|
| Anatomy & Details | Highly stable finger and limb structures with a massive reduction in ghosting/artifacts. |
| Physical Motion | Smooth, fluid transitions adhering naturally to inertia and gravity. |
| Shot Transitions | Flawless cinematic cutting logic with precise timestamp orchestration. |
| Visual Aesthetics | Composition, lighting, and overall cinematic atmospheric depth are heavily enhanced. |
| Character Consistency | Natural facial expressions even in tight close-ups; prevents sudden "face-swapping." |
| High Dynamic Limits | Meets the vast majority of movement demands. Slight motion blur may still occur during extreme, highly complex actions. This is currently being addressed via optimized post-processing workflows—stay tuned! |
⚙️ Usage Guide
- Recommended Base Model:
ltx-2.3-22b-distilled-1.1_transformer_only_fp8_scaled.safetensors - ComfyUI Workflow: Uploaded and available in the files tab of this repository. Highly recommended to use in First & Last Frame Mode for ultimate scene control.
- Online Demo: Click here to try it online
📝 Exclusive: Singularity Prompting Framework
This model follows a strict prompt structure to unlock its full cinematic potential. Please adhere closely to the "Cinematic Timeline Structure" below.
💡 Core Rule: Keep visual descriptions, timestamps, actions, and dialogue strictly formatted in English as shown below.
📐 Output Template Structure
[Scene & Style]: Core visual description in one sentence (e.g., Cinematic wuxia style, dim lighting, Anime, 3D).
[Action Timeline]: 0-X seconds, [action / emotional description].
[Camera Timeline]: 0-X seconds, [camera movement / composition parameters].
[Environment]: Lighting source, contrast, and color grading details.
[Dialogue]: 0-X seconds, [Character] says: "[Dialogue text]".
[Audio & Technical]: Background sounds, film grain, subtitle exclusion commands, etc.
---
#### 🎬 Example Prompt
Cinematic wuxia style, indoor dim lighting, mysterious mood. 0-10 seconds, young man in ancient white robes looks down with a confused expression. 0-10 seconds, tight close-up, static camera with slight handheld movement. Dark stone background, warm candlelight bokeh. 0-10 seconds, man says: "What on earth is this? I've never heard of it before.". Voice: low and confused, Pace: slow. Precise lip-sync, film grain, cinematic bokeh, no subtitles.
---
# Contact Information
WeChat: aigctyd
Email: a592991299@gmail.com