Model Card: UltraVAD UltraVAD is a context aware, audio native endpointing model. It estimates the probability that a speaker has finished their turn in real time by fusing recent dialog text with the user’s audio. UltraVAD consumes the dialogue history and the last user audio turn, then produces a probability for the end of turn token . Model Details Developer: Ultravox.ai Type: Context aware audio–text fusion endpointing Backbone: Llama 8B (post trained) Languages (26): ar, bg, zh, cs, da, nl, en, fi, fr, de, el, hi, hu, it, ja, pl, pt, ro, ru, sk, es, sv, ta, tr, uk, vi What it predicts. UltraVAD computes the probability P( context, user audio) Sources Website/Repo: https://ultravox.ai Demo: https://demo.ultravox.ai/ Benchmark: https://huggingface.co/datasets/fixie ai/turntaking contextual tts Blog: https://www.ultravox.ai/blog/ultravad is now open source introducing the first context aware audio native endpointing model Usage Use UltraVAD as a turn taking oracle in voice agents. Run it alongside a lightweight streaming VAD; when short silences are detected, call UltraVAD and trigger your agent’s response once the probability crosses your threshold. Training Text only post train…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy