Skip to content
Cutshed

Audio · 27

AI audio maker

The Audio tab in Studio writes a song, reads text aloud, or builds a sound effect. The free audio tools on this site still keep the file on your device. This one sends the prompt to the model.

Music, speech, and sound effects. A finished file plays in Studio with a waveform, then downloads as MP3 or WAV, or opens in the editor.

Open Audio in Studio

See Studio pricing

How to use it

Open Studio and choose Audio, next to Emoji maker. Music takes a prompt, a genre, and lyrics when the model sings. Lyria 2 is a fixed 30-second clip. MiniMax songs are a full track.

Speech is ElevenLabs Turbo, MiniMax Speech HD, or Kokoro. Pick a voice, set speed, and set stability when the model has it. The character counter shows the credit cost before you run it. Kokoro Heart has a published preview.

Sound effects take a description and a length from 2 to 22 seconds. Finished audio plays in your history with a waveform. Download MP3 or WAV, delete the job, or send it to the editor.

Tips

  • Kokoro is the cheap voice and the weights are Apache-2.0, so the speech can go in a paid video. ElevenLabs and MiniMax are paid routes with commercial use, not a film or TV clearance.
  • MusicGen is not in the list. Its weights are non-commercial. ElevenLabs Music is not in the list either: on fal it is $0.80 a minute.
  • A failed or blocked run puts the credits back. Voice cloning is not part of this tab.

FAQ

How many credits is a song?

MiniMax Music 2.0 is 5 credits. Google Lyria 2 is 15 credits for 30 seconds. MiniMax Music 2.5 is 25 credits. A sound effect is 5 credits up to 15 seconds and 10 credits at 22 seconds.

How is speech priced?

Kokoro is about $0.02 per 1,000 characters, ElevenLabs Turbo $0.05, and MiniMax Speech HD $0.10. Credits are that cost times 1.5, rounded up to the next 5. One thousand characters is 5, 10, or 15 credits.

Can I use the audio commercially?

Kokoro output can. MiniMax music and speech can, through fal, as long as you do not copy a copyrighted melody or a living artist's voice. ElevenLabs speech and sound effects can on a paid API, with the limits of that plan. This page does not add voice cloning.

License

Music uses MiniMax Music 2.0, MiniMax Music 2.5, and Google Lyria 2 on fal. Speech uses ElevenLabs Turbo v2.5, MiniMax Speech-02 HD, and Kokoro (Apache-2.0). Sound effects use ElevenLabs Sound Effects v2. Meta MusicGen is excluded because the weights are CC-BY-NC.