This repository contains the SECOND ITERATION mHuBERT 147 model. The best mHuBERT 147 model is available here. MODEL DETAILS: 2nd iteration, K=1000, HuBERT base architecture (95M parameters), 147 languages. Table of Contents: 1. Summary 2. Training Data and Code 3. ML SUPERB Scores 4. Languages and Datasets 6. Citing and Funding Information mHuBERT 147 models mHuBERT 147 are compact and competitive multilingual HuBERT models trained on 90K hours of open license data in 147 languages. Different from traditional HuBERTs, mHuBERT 147 models are trained using faiss IVF discrete speech units. Training employs a two level language, data source up sampling during training. See more information in our paper. This repository contains: Fairseq checkpoint (original); HuggingFace checkpoint (conversion using transformers library); Faiss index for continuous pre training (OPQ16 64,IVF1000 HNSW32,PQ16x4fsr). Related Models: 3rd Iteration mHuBERT 147 (best) 1st Iteration mHuBERT 147 HUTTER 12 CommonVoice Prototype (12 languages) Training Manifest list available here. Please note that since training, there were CommonVoice removal requests. This means that some of the listed files are no longer av…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy