Cancel the center of a stereo mix, where the lead vocal usually lives, and download the instrumental as a WAV.
It makes a karaoke version of a stereo song by removing what is common to both channels. A lead vocal is normally panned dead center, so it appears identically on the left and the right, and subtracting one channel from the other makes it disappear while anything panned to a side survives.
The whole method is a few lines of arithmetic per sample, and it explains every strength and weakness at once.
This is phase cancellation, not source separation, so it removes a position in the stereo field rather than a singer. Anything else parked there leaves with the vocal: snare, bass, a centered lead synth. Anything the vocal spills into stays: stereo reverb tails, delays, doubled takes panned wide. Modern mastering makes it harder still, since wideners and multiband processing shift a voice slightly off center in a way no fixed subtraction can follow. Mono input cannot be processed, and the tool says so instead of returning silence.
No. The file is decoded in this tab and processed in a Web Worker on your own machine, and the WAV is generated locally. Turn off your network before you drop the file in and it still works.
This method works on the difference between the left and right channels. A mono file has no difference, so subtracting one from the other leaves silence. Any tool that claims otherwise is doing something else, usually a spectral model. Find a stereo copy.
Two reasons. First, the low end rescue: to keep the kick and bass, the tool adds a gently filtered copy of the center back in, and the lowest part of the voice returns with it, roughly 13 dB down around 660 Hz. Second, the voice is not perfectly centered. Reverb and delay on a vocal are almost always stereo, so the wet tail survives even when the dry voice cancels. Doubled or harmonised vocals are usually panned apart and survive too. And a modern master applies stereo widening and multiband compression that move a voice off dead center in ways this cannot follow.
Because every other center panned element goes with the vocal. Kick, snare and bass usually sit dead center, so they cancel too. The low end rescue brings back the kick and the bass fundamentals, but the snare body and any centered lead instrument are still casualties, and what remains is the sides, so the stereo image goes too.
No, and the difference is worth being honest about. Model based separators such as Demucs are trained on thousands of songs and can isolate a voice that is not centered, but they need a GPU server, which means uploading your file and waiting. This is deterministic arithmetic on your own machine that finishes in about a second per song. For well produced pop with a dry centered lead it does a surprisingly good job.
Sing over it, practise over it, or drop the WAV into Gusta Music, the free browser studio this tool comes from, and build a new arrangement around it. Removing a vocal gives you no rights to the recording, so keep commercial use to material you own or have licensed.
Gusta Music is the free browser studio these tools came out of: multitrack timeline, sampling, mic recording, and an AI copilot that edits the song with you instead of generating one for you. Your work stays yours.
Open the studio