Model Card for Finetuned Version of Whisper Medium This model was trained on a subset of the synthetically generated data that later on was filtered to increase the performance of Whisper Model. The approach involves aligning representations of synthetic audio and corresponding text transcripts to identify and remove low quality samples, improving the overall training data quality. In this Specific Model we used 96,08% of synthetic data generated by SeamllesMT4LargeV2, the rest was removed by the filtering model. The training set also contained, the CommonVoice Dataset, Multilibri Speach, and Bracarense (Fully Portuguese Dialect) Model Details Developed by: Yuriy Perezhohin, Tiago Santos, Victor Costa, Fernando Peres, and Mauro Castelli. Funded by: Remynd Shared by: Remynd Model type: ASR with contrastive learning based synthetic data filtering Language: Portuguese License: APACHE 2.0 Finetuned from model: Whisper Small Model Sources Repository: https://github.com/my north ai/semantic audio filtering Paper: (https://ieeexplore.ieee.org/document/10720758/) Uses This model can be directly used for improving ASR systems in Portuguese, particularly in scenarios with limited real world…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy