Model Card for EnCodec This model card provides details and information about EnCodec 32kHz, a state of the art real time audio codec developed by Meta AI. This EnCodec checkpoint was trained specifically as part of the MusicGen project, and is intended to be used in conjuction with the MusicGen models. Model Details Model Description EnCodec is a high fidelity audio codec leveraging neural networks. It introduces a streaming encoder decoder architecture with quantized latent space, trained in an end to end fashion. The model simplifies and speeds up training using a single multiscale spectrogram adversary that efficiently reduces artifacts and produces high quality samples. It also includes a novel loss balancer mechanism that stabilizes training by decoupling the choice of hyperparameters from the typical scale of the loss. Additionally, lightweight Transformer models are used to further compress the obtained representation while maintaining real time performance. This variant of EnCodec is trained on 20k of music data, consisting of an internal dataset of 10K high quality music tracks, and on the ShutterStock and Pond5 music datasets. Developed by: Meta AI Model type: Audio Code…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy