Huggingface Implementation of AV HuBERT on the MuAViC Dataset This repository contains a Huggingface implementation of the AV HuBERT (Audio Visual Hidden Unit BERT) model, specifically trained and tested on the MuAViC (Multilingual Audio Visual Corpus) dataset. AV HuBERT is a self supervised model designed for audio visual speech recognition, leveraging both audio and visual modalities to achieve robust performance, especially in noisy environments. Key features of this repository include: Pre trained Models: Access pre trained AV HuBERT models fine tuned on the MuAViC dataset. The pre trained model been exported from MuAViC repository. Inference scripts: Easily pipelines using Huggingface’s interface. Data preprocessing scripts: Including normalize frame rate, extract lips and audio. Inference code Data preprocessing scripts Pretrained AVSR model Languages Huggingface Arabic Checkpoint AR German Checkpoint DE Greek Checkpoint EL English Checkpoint EN Spanish Checkpoint ES French Checkpoint FR Italian Checkpoint IT Portuguese Checkpoint PT Russian Checkpoint RU Multilingual Checkpoint ar de el es fr it pt ru Acknowledgments AV HuBERT : A significant portion of the codebase in this…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy