AudioLDM 2 Large AudioLDM 2 is a latent text to audio diffusion model capable of generating realistic audio samples given any text input. It is available in the 🧨 Diffusers library from v0.21.0 onwards. Model Details AudioLDM 2 was proposed in the paper AudioLDM 2: Learning Holistic Audio Generation with Self supervised Pretraining by Haohe Liu et al. AudioLDM takes a text prompt as input and predicts the corresponding audio. It can generate text conditional sound effects, human speech and music. Checkpoint Details This is the original, large version of the AudioLDM 2 model, also referred to as audioldm2 full large 1150k . There are three official AudioLDM 2 checkpoints. Two of these checkpoints are applicable to the general task of text to audio generation. The third checkpoint is trained exclusively on text to music generation. All checkpoints share the same model size for the text encoders and VAE. They differ in the size and depth of the UNet. See table below for details on the three official checkpoints: Checkpoint Task UNet Model Size Total Model Size Training Data / h audioldm2 Text to audio 350M 1.1B 1150k audioldm2 large Text to audio 750M 1.5B 1150k audioldm2 music Text…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy