At a glance

Accepts
MP3, WAV, M4A, AAC, OGG, Opus, FLAC, and the audio track of a video
Outputs
Two 44.1 kHz stereo WAV files, one voice and one music
Where it runs
Entirely in your browser, on your device
What leaves your device
Nothing - the audio never leaves your device; the accurate mode downloads model weights to you
Limits
Accurate mode needs a one-time 166 MB download and is slow without WebGPU; the model is trained on sung vocals, so spoken word is out of distribution; the fast mode only works on a steady musical bed
Account
Not required - there is no sign-up
Price
Free, with no watermark and no usage cap

Splitting a voice out of background music is not a filtering problem, even though it sounds like one. Speech and music occupy the same part of the spectrum: a speaking voice has its fundamental around 85 to 255 Hz, most of what makes it intelligible sits between 300 Hz and 4 kHz, and its consonants reach past 10 kHz. Guitars, piano, strings and snare live in exactly that band. Any filter wide enough to keep the voice keeps most of the instruments with it.

The two approaches

How to use it

  1. Drop in an audio file, or a video to pull its audio track.
  2. Pick fast or accurate. Try fast first; it costs nothing and tells you quickly whether your file is an easy case.
  3. Play both results in the page, then download either as WAV.

What it will not do

The model was trained on sung vocals, so spoken word over music is slightly outside its training. It usually still works on dialogue and podcast beds, but that is the case worth testing rather than assuming. Neither method can invent detail that was masked in the original recording, so a voice buried under a loud mix comes back with audible artefacts. And the music track is derived by subtracting the voice from your original, which means it inherits whatever the voice track got wrong.

Frequently asked questions

Is my audio uploaded to a server?

No. The audio is decoded and separated in your browser tab. The only network request this tool makes is downloading the separation model itself, which is a download, not an upload - your file never leaves your device.

Can you separate voice from music just by filtering frequencies?

Not really. Speech sits at roughly 85-255 Hz for the fundamental with intelligibility up to about 4 kHz and consonants to 10 kHz, and music covers all of that same range. A filter that keeps the voice keeps most of the instruments too. Removing the very low bass and the very high cymbal air helps a little, but it is cleanup rather than separation.

What does the fast mode actually do?

It estimates the steady musical background from the median level of each frequency over time, then subtracts it. A musical bed holds the same spectral shape for minutes while a voice moves around, so the median is mostly the bed. It works on a steady backing track and falls apart when the music changes as much as the voice does.

Why is the accurate mode a 166 MB download?

It is a real neural network - HT-Demucs, fine-tuned for vocals, exported to ONNX. It runs entirely in your browser, so the weights have to come to you. It is cached after the first time, so later runs start immediately.

Does it work on speech, or only singing?

The model is trained on sung vocals, so spoken word over music is slightly outside what it learned. In practice it usually still separates dialogue and podcast beds well, but it is the one case where it is worth trying the file before relying on it.

Why do the two output files sum back to my original?

The music track is calculated as your original minus the isolated voice, rather than separately. That means nothing is lost in the gap between the two, and you can layer them back to get exactly what you started with.

What formats can I put in?

Anything your browser can decode: MP3, WAV, M4A, AAC, OGG, Opus, FLAC, and the audio track of a video file. Output is 44.1 kHz stereo WAV.