Opus Sounds Directory
Blog

Open-Source AI Audio Models on Hugging Face: Music, Sound Effects and Video-to-Audio

A guide to open AI audio models on Hugging Face (MusicGen, AudioGen, Stable Audio Open, AudioLDM 2, ACE-Step, MMAudio) with what each does and its license.

The main open AI audio models on Hugging Face are MusicGen and ACE-Step for music, AudioGen and AudioLDM 2 for sound effects, Stable Audio Open for short music and sounds, and MMAudio and HunyuanVideo-Foley for adding sound to video. Their licenses differ a lot, and several are non-commercial, so read the model card before using output in client work.

Here is what each one does, how long a clip it makes, and what its license says.

A text-to-audio model is a generative model that takes a written description, such as "rain on a tin roof" or "upbeat synth pop, 120 BPM", and outputs an audio waveform.

A video-to-audio model takes video frames (often plus a text hint) and generates sound that is synchronized to what happens on screen.

You can browse both on Hugging Face's text-to-audio model list.

MusicGen (Meta, AudioCraft). A text-to-music model available in several sizes, including musicgen-small, musicgen-large and a melody-conditioned version. It runs in the Transformers library (MusicGen docs) and in Meta's AudioCraft repo. AudioCraft's code is MIT, but the model weights are CC-BY-NC 4.0 (non-commercial). Try it in the MusicGen Space.

Stable Audio Open (Stability AI). Stable Audio Open 1.0 generates variable-length stereo audio up to 47 seconds at 44.1 kHz, and Stable Audio Open Small goes up to 11 seconds. Both use the Stability AI Community License, and you accept terms on the model page to download. It runs with stable-audio-tools or the Diffusers Stable Audio pipeline.

ACE-Step. ACE-Step 1.5 is a music generation model released under the MIT license. Its model card says generated music can be used commercially and that it scales from short loops to long compositions. The earlier ACE-Step v1 3.5B is Apache-2.0.

AudioGen (Meta). audiogen-medium is trained for text-to-sound effects, with generation length set in code (the card's example uses 5 seconds). Weights are CC-BY-NC 4.0.

AudioLDM 2. AudioLDM 2 is a latent diffusion model for text-conditional sound effects, speech and music, available through the Diffusers AudioLDM 2 pipeline. License: CC-BY-NC-SA 4.0.

Bark (Suno). Bark is mainly a text-to-speech model, but its card says it can also produce music, background noise, simple sound effects and nonverbal sounds like laughing. License: MIT.

MMAudio. MMAudio generates audio from video (and text). The code is MIT, and the checkpoints are CC-BY-NC 4.0 for non-commercial use.

HunyuanVideo-Foley (Tencent). HunyuanVideo-Foley generates 48 kHz sound effects synchronized to video, guided by a text description. It uses the Tencent Hunyuan community license.

Model Main use License on the model card
MusicGen Music CC-BY-NC 4.0 (weights)
AudioGen Sound effects CC-BY-NC 4.0
Stable Audio Open 1.0 / Small Music and sounds, up to 47 s / 11 s Stability AI Community License
AudioLDM 2 Sound effects, music, speech CC-BY-NC-SA 4.0
ACE-Step 1.5 Music MIT
Bark Speech, some sounds MIT
MMAudio Video-to-audio CC-BY-NC 4.0 (checkpoints)
HunyuanVideo-Foley Video-to-audio Tencent Hunyuan community license

CC-BY-NC is a Creative Commons license that allows reuse with credit but forbids commercial use. Licenses and model versions change, so treat this table as a starting point and read each card yourself.

Generative models are good at texture and realism. They are weaker at exact length, frame-accurate hits and identical re-renders. When you need a whoosh on frame 30 or a bed that is exactly 720,000 samples, it's often easier to have an LLM write synthesis code.

Opus Sounds Directory collects prompts in that style next to the Python code, spectrogram and measured loudness. Each entry is Claude Opus 5.5 output: the model writes the synthesis code, and the page names the model. Here are two examples and their prompts.

Make a 10 s music bed for a vertical video ad at 30 fps.
Output: one WAV, 48 kHz / 16-bit / stereo, exactly 480000 samples.
Python + numpy/scipy only. Synthesise everything (no samples, no downloads). Fixed random seed.
Timing: 100 BPM (1 beat = 18 frames). Cues: frame 0 = alert stack; frame 120 = notifications mute;
last beat = soft ending that fades to exactly 0.
Sound: chaos (stacked notification pings, rapid clicks, tense pulse) -> calm (airy pad,
slow exhale noise bed, single soft bell), key F major.
Master: -14 LUFS integrated, true peak <= -1 dBTP.
Verify length, loudness and peak with ffmpeg ebur128; confirm the energy drop at frame 120.
Make a 0.4 s UI error sound at 30 fps.
Output: one WAV, 48 kHz / 16-bit / stereo, exactly 19200 samples.
Python + numpy/scipy only. Synthesise everything. Fixed random seed.
Sound: low muted thud (sine at 180 Hz, 3 ms attack, 150 ms decay) plus a quiet minor-second interval
(E5 + F5, 60 ms). Transient at frame 0, nothing above 6 kHz, 5 ms fade-out to exactly 0.
Master: true peak <= -3 dBTP. Verify length and peak with ffmpeg ebur128 and report the numbers.

A good hybrid workflow: generate texture with a model (rain, crowd, a music idea), then use code for the timed parts (hits, risers, stings) and mix both.

Describe the sound plainly and add the facts the model can use:

  • Genre or sound source ("lo-fi hip hop", "heavy wooden door").
  • Tempo and key for music, if the model responds to them.
  • Instrumentation or texture ("muted piano, vinyl crackle").
  • Duration, set in the code or UI rather than the text where possible.
  • What to avoid ("no vocals", "no reverb tail").

Then generate several takes and pick by ear. These models aren't deterministic unless you fix the seed.

FAQ

What is the best open-source AI music model?
It depends on your use. MusicGen and Stable Audio Open are widely used for short clips, and ACE-Step 1.5 targets full songs under an MIT license; test a few against your own prompts.
Can I use music from MusicGen commercially?
Be careful. Meta releases the MusicGen weights under CC-BY-NC 4.0, a non-commercial license, so check the model card and get legal advice before commercial use.
Which open model makes sound effects rather than music?
AudioGen is trained for text-to-sound effects, AudioLDM 2 covers general sound effects and music, and Stable Audio Open generates both short sounds and music.
Is there an open model that adds sound to a video?
Yes. MMAudio and HunyuanVideo-Foley are video-to-audio models on Hugging Face that generate sound effects synchronized to video frames.
Do I need a GPU to run these models?
For reasonable speed, usually yes, but many models have hosted demos in Hugging Face Spaces that you can try in a browser first.