Universal Audio Annotation Pipeline — self contained model mirror This repository is a complete, self contained mirror of the LAION Universal Audio Annotation Pipeline : all model weights for every stage plus all the code needed to run it. If any of the upstream model repositories ever disappears, cloning this single repo gives you everything required to reproduce the pipeline end to end. 💻 Code (GitHub): https://github.com/LAION AI/univeral audio annotation pipeline 🌐 Live example predictions: https://laion ai.github.io/univeral audio annotation pipeline/predictions/ (also bundled here under predictions/ ) What it does Given any audio file (a movie scene, a podcast, a field recording…), the pipeline produces a single structured JSON annotation of everything audible , second by second across the whole clip: Speech — transcription, speaker diarization (who speaks when), language, accent, speaking rate, age, gender, voice timbre, and expressive captions for the speaker's emotion and speaking style (e.g. "clearly intense anger laced with wounded disappointment" , "low conspiratorial whisper" ); singing is flagged explicitly. Vocal bursts — laughs, gasps, sighs, screams, scoffs, etc.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy