Vocal Remover for Songs Online

AI-powered stem separation — extract instrumentals, isolate acapella, or create karaoke versions from any song. Remove vocals using machine learning for music production, practice, and remix.

How to Remove Vocals from a Song

1

Upload your audio file (MP3, WAV, FLAC, or any other format)

2

Wait a few seconds while the tool separates vocals from music

3

Preview the result — listen to Music, Vocal, or both together

4

Adjust the Separation Strength slider for optimal results

5

Choose your output format (MP3, WAV, or FLAC)

6

Download the instrumental (Music) or vocal track separately

Get instrumental and vocal tracks from any song in seconds

Drop your audio file here
MP3, WAV, OGG, FLAC, AAC, M4A, WEBM, AIFF, OPUS, WMA, M4B, CAF
How to separate
Processing variants
This file is mono — Aggressive, Two-pass and Ultra need a stereo file (the separation algorithms rely on the stereo image to isolate the vocal). Standard will still work.

Listen and adjust

Music
Vocal
0:00 0:00
Quick mix:
Separation Strength 90%

Save

Advanced export options — per-stem formats, fades, LUFS normalization
Loudness normalization Apply EBU R128 loudnorm filter to match streaming/broadcast LUFS targets. Adds ~3-8 s to encoding.

Every track can be saved on its own — the button in its row. “Mix” saves what you hear: current volumes and the selected track.

💡 How this works: In “Ultra” a neural model does the separation — it recognizes the voice by timbre and does not depend on where the vocal sits in the stereo image. “Fast” works differently: it subtracts whatever sounds identical in the left and right channels — near-instant, but rougher. The Separation strength slider balances the two tracks.

What runs under the hood

The modes are not “quality levels” of one algorithm — they are different engines. Two you pick yourself, two switch on automatically. “Ultra” downloads a model first, which is worth knowing on mobile data.

Mode What runs Download
Fast · lower quality Stereo-image separation: what sits in the center is subtracted spectrally (4096-sample window, voice band 100 Hz – 12 kHz). Nothing — starts immediately
AI A neural separation model (ONNX). Not selectable — it switches on by itself for mono files, where “Ultra” cannot run. About 56 MB, once
Ultra · best quality MDX-Net Kim Vocal 2 — the most accurate available and the slowest. Needs stereo. About 64 MB, once
Backing vocal UVR MDX-Net Karaoke 2: a second pass over the finished vocal that splits lead from backing. About 50 MB, once

What to expect

  • “Fast” needs stereo: a mono recording has no stereo image to subtract, so both tracks come back identical. “Ultra” needs stereo as well and does not run on mono — there a different neural engine takes over automatically.
  • A model is downloaded once and stays in the browser cache — the next song in the same browser starts straight away.
  • Perfect separation does not exist: the denser the mix, the more often a trace of the voice stays in the backing track. The tool measures that residue itself and offers a stronger variant.
Published Updated Author: