A High Quality and Large Scale Dataset for English Vietnamese Speech Translation PhoST is a high quality and large scale English Vietnamese speech translation dataset with 508 audio hours, consisting of 331K triplets of (sentence lengthed audio, English source transcript sentence, and Vietnamese target subtitle sentence). Details of the dataset construction and experimental results can be found in our INTERSPEECH 2022 paper: @inproceedings{PhoST, title = {{A High Quality and Large Scale Dataset for English Vietnamese Speech Translation}}, author = {Linh The Nguyen and Nguyen Luong Tran and Long Doan and Manh Luong and Dat Quoc Nguyen}, booktitle = {Proceedings of the 23rd Annual Conference of the International Speech Communication Association (INTERSPEECH)}, year = {2022} } By downloading this dataset, USER agrees: to use the dataset for research or educational purposes only. to not distribute the dataset or part of the dataset in any original or modified form. and to cite our INTERSPEECH 2022 paper "A High Quality and Large Scale Dataset for English Vietnamese Speech Translation" whenever the dataset is used to help produce published results. For further information or requests, p…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy