Stable Audio Open 1.0: Open Source Text-to-Audio Generation by Stability AI

ComfyUI Wiki

Stable Audio Open 1.0 is Stability AI's open-source text-to-audio model capable of generating ambient music, foley, and sound effects up to 47 seconds.

S

Stable Audio Open 1.0

Text-to-AudioSound EffectsFoleyOpen Source

Open-source text-to-audio model by Stability AI. Generates ambient music, foley, and sound effects up to 47 seconds from natural language prompts. Built on a latent diffusion architecture with a 1.2B parameter U-Net and a text conditioning module.

DeveloperStability AI
Release Date2024-08
Architecture1.2B Latent Diffusion U-Net
LicenseStability-AI-Open-RAIL-M
Max Duration47 seconds
Sample Rate44.1 kHz
CapabilitiesAmbient music, foley, sound effects

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…