Stable Audio 3 Medium (Base) Note: This is the base (pre trained) model intended for fine tuning. If you are looking to generate audio directly, please use Stable Audio 3 Medium instead. Please note: For commercial use, please refer to https://stability.ai/license Model Description Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable length audio generation and editing. Since our models can generate several minutes of audio, variable length generations are key to avoid the cost of producing full length generations for short sounds. We also support inpainting, enabling targeted audio editing and the continuation of short recordings. Our latent diffusion models operate on top of a novel semantic acoustic autoencoder that projects audio into a compact latent space, enabling efficient diffusion based generation while preserving audio fidelity and encouraging semantic structure in the latent. Finally, we run adversarial post training to both accelerate inference and improve generation quality, reducing the number of inference steps while improving fidelity and prompt adherence. Stable Audio 3 models are trained on licensed and Creative Commons d…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy