Vocal Remover for Songs Online
AI-powered stem separation — extract instrumentals, isolate acapella, or create karaoke versions from any song. Remove vocals using machine learning for music production, practice, and remix.
How to Remove Vocals from a Song
Upload your audio file (MP3, WAV, FLAC, or any other format)
Wait a few seconds while the tool separates vocals from music
Preview the result — listen to Music, Vocal, or both together
Adjust the Separation Strength slider for optimal results
Choose your output format (MP3, WAV, or FLAC)
Download the instrumental (Music) or vocal track separately
Get instrumental and vocal tracks from any song in seconds
Listen and adjust
Save
Advanced export options — per-stem formats, fades, LUFS normalization
Every track can be saved on its own — the button in its row. “Mix” saves what you hear: current volumes and the selected track.
Custom variant settings
Power-user controls — adjust each stage of the pipeline. Saved settings persist across sessions.
What runs under the hood
The modes are not “quality levels” of one algorithm — they are different engines. Two you pick yourself, two switch on automatically. “Ultra” downloads a model first, which is worth knowing on mobile data.
| Mode | What runs | Download |
|---|---|---|
| Fast · lower quality | Stereo-image separation: what sits in the center is subtracted spectrally (4096-sample window, voice band 100 Hz – 12 kHz). | Nothing — starts immediately |
| AI | A neural separation model (ONNX). Not selectable — it switches on by itself for mono files, where “Ultra” cannot run. | About 56 MB, once |
| Ultra · best quality | MDX-Net Kim Vocal 2 — the most accurate available and the slowest. Needs stereo. | About 64 MB, once |
| Backing vocal | UVR MDX-Net Karaoke 2: a second pass over the finished vocal that splits lead from backing. | About 50 MB, once |
What to expect
- “Fast” needs stereo: a mono recording has no stereo image to subtract, so both tracks come back identical. “Ultra” needs stereo as well and does not run on mono — there a different neural engine takes over automatically.
- A model is downloaded once and stays in the browser cache — the next song in the same browser starts straight away.
- Perfect separation does not exist: the denser the mix, the more often a trace of the voice stays in the backing track. The tool measures that residue itself and offers a stronger variant.