Audio · 01
Vocal remover
Turn a stereo mix into a vocal WAV and an instrumental WAV. Open-Unmix UMX-HQ runs in the tab, in sections, so a longer song does not have to sit in one giant spectrogram.
How to get a usable instrumental
Start from the least-processed stereo file you have. A WAV, FLAC, or high-bitrate AAC keeps more of the vocal’s overtones than a 64 kbps MP3, and the model is less likely to smear the vocal into the accompaniment. Mono files are duplicated to stereo so the network can run, but a true stereo mix separates better.
The network was trained as Open-Unmix on MUSDB18-HQ. It reads a 44.1 kHz magnitude spectrogram (4096-point Hann window, hop 1024) and estimates the vocal magnitude. This page applies that estimate as a ratio mask on the mix phase, inverse-transforms the vocal, and builds the instrumental by subtracting the vocal from the original samples. That residual is what you want for a karaoke bed: whatever the model did not call vocal stays in the track, including the high end above the model’s useful band.
Songs are cut into roughly eight-second sections with a little context on either side, because the network is a bidirectional LSTM and a full-length spectrogram will exhaust a laptop tab. Overlaps are crossfaded. Eight minutes or 80 MB is the hard stop. Phones are happier under three minutes.
Preview both players on headphones before you lay the instrumental under a voiceover. If the chorus vocal is still sitting in the bed, the take is probably wet or doubled. You will not EQ that out inside this tool. Bounce the WAV and finish it in your editor.
Tips
- Trim to the section you actually need. A thirty-second hook is a much kinder job than a full album track on a phone.
- If the source is a video, extract the audio with Video to MP3 first, or drop the video here. The picture is ignored.
- Listen to the vocal stem against the lyric. Missing words usually mean the vocal was never isolated, not that the WAV encoder dropped them.
- Keep the WAV if you will edit further. Make an MP3 only when you need a smaller file to send.
FAQ
Does the song get uploaded?
No. The browser downloads the separation model, then runs it locally. The audio samples stay in the tab’s memory and are discarded when you leave or pick another file.
Why is there a download before the first song?
The Open-Unmix vocals network is about 17 MB and is not packed into the page script. ONNX Runtime adds roughly 14 MB for WASM or 27 MB if this browser can use WebGPU. Both are cached afterwards.
Will it use WebGPU or WASM?
The tool asks for a WebGPU adapter first and runs one frame as a check. Open-Unmix casts magnitudes to float16, so an adapter without shader-f16 fails that check and the same job continues on WASM. The status line names which one finished.
Why only vocals and an instrumental?
This version estimates the vocal with a stereo Open-Unmix model and builds the instrumental by subtracting that vocal from the mix. Drums, bass, and other are not separate files.
Can I use the result in a commercial video?
The model weights are MIT licensed, so you may use the software commercially. That license does not cover the recording. You still need the rights, a license, or a release for the song itself.
Why is the vocal still audible in the instrumental?
Reverb, doubles, and live-room bleed share spectrum with the vocal. Dry, centered studio vocals separate more cleanly than a phone recording of a band in a room.
Model and license
Weights are Open-Unmix UMX-HQ vocals, MIT, published by Inria / SigSep and exported to ONNX by edgetools/umx-hq (17,820,856 bytes, SHA-256 3d05709ff7197bbd4a33aee759e4e82002acb491f41c94770227ab57507b6ccb). UMX-L was not used because those weights are CC BY-NC-SA and not cleared for commercial use. Cite Stöter, Uhlich, Liutkus, and Mitsufuji, Journal of Open Source Software, 2019, doi:10.21105/joss.01667.