VoxCommunis Corpus The VoxCommunis Corpus is a phonetic corpus derived from the Mozilla Common Voice Corpus. Corresponding audio files and corpus metadata can be downloaded from Mozilla Common Voice, or from one of several Hugging Face repositories for the differing versions. Within each folder, the filenames share similar structure and contain critical information for effectively using the file. More detail regarding the specifics of the filename for each file type is provided below. In general, a filename such as mk xpf lexicon19 corresponds to: + mk : Common Voice language ID code (Macedonian) + xpf : G2P system for the lexicon (XPF Corpus) + 19 : Common Voice version (Macedonian Version 19) acoustic models/ : The acoustic models have been trained using the Montreal Forced Aligner, and the force aligned TextGrids are obtained directly from those alignments. These acoustic models can be downloaded and re used with the Montreal Forced Aligner for new data. lexicons/ : The lexicons are developed using various toolkits. Some manual correction has been applied, and we hope to continue improving these. Any updates from the community are welcome. + epi : Epitran + xpf : XPF Corpus + ch…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy