TADA: A Generative Framework for Speech Modeling via Text Acoustic Dual Alignment A unified speech language model that synchronizes speech and text into a single, cohesive stream via 1:1 alignment. Text Acoustic Dual Alignment Large Language Model TADA is a unified speech language model that synchronizes speech and text into a single, cohesive stream via 1:1 alignment. By leveraging a novel tokenizer and architectural design, TADA achieves high fidelity synthesis and generation with a fraction of the computational overhead required by traditional models. ⭐️ arxiv: https://arxiv.org/abs/2602.23068 \ ⭐️ demo1: https://huggingface.co/spaces/fffiloni/tada dual alignment tts demo \ ⭐️ demo2: https://huggingface.co/spaces/HumeAI/tada \ ⭐️ github: https://github.com/HumeAI/tada \ ⭐️ blog post: https://www.hume.ai/blog/opensource tada \ Key Features 1:1 Token Alignment: Unlike standard models, TADA’s tokenizer encodes audio into a sequence of vectors that perfectly matches the number of text tokens. Dynamic Duration Synthesis: As a TTS model, it generates the full speech segment for a text token in a single autoregressive step, regardless of length. This eliminates the need for fixed frame…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy