Vocos: Closing the gap between time domain and Fourier based neural vocoders for high quality audio synthesis Audio samples Paper [[abs]](https://arxiv.org/abs/2306.00814) [[pdf]](https://arxiv.org/pdf/2306.00814.pdf) Vocos is a fast neural vocoder designed to synthesize audio waveforms from acoustic features. Trained using a Generative Adversarial Network (GAN) objective, Vocos can generate waveforms in a single forward pass. Unlike other typical GAN based vocoders, Vocos does not model audio samples in the time domain. Instead, it generates spectral coefficients, facilitating rapid audio reconstruction through inverse Fourier transform. Installation To use Vocos only in inference mode, install it using: If you wish to train the model, install it with additional dependencies: Usage Reconstruct audio from mel spectrogram Copy synthesis from a file: Citation If this code contributes to your research, please cite our work: License The code in this repository is released under the MIT license.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy