This is one of those topics where the received wisdom is half right, which is worse than being wrong.
AI separation has moved from a novelty to a tool people ship records with. It is worth understanding what it estimates, because that tells you exactly where it will let you down.
Nothing below is theory for its own sake. Every number here is one you can check on a meter, and every step is one you can run on the next track you open.
Key takeaways
- Separation estimates sources, it does not extract hidden multitracks.
- Bass and drums separate most cleanly, piano and guitar least.
- Feed it lossless audio, because lossy sources lose the detail the model needs.
- Heavily limited masters separate worse than dynamic ones.
How AI stem separation actually works
A separation model is trained on many thousands of songs where the individual parts were already known. It learns what a vocal, a snare, a bass line and a piano look like in the time-frequency domain, then estimates a mask for each source in a mixed signal it has never heard.
It is estimation, not extraction. The original multitrack is not hidden inside the stereo file waiting to be recovered. The model is making an informed reconstruction, and that is why results vary with the material.
For vocals: The hardest source to isolate cleanly because the human voice overlaps almost every other instrument in the 200 Hz to 5 kHz range.
What separates cleanly, and what does not
| Source | Typical result | Why |
|---|---|---|
| Bass | Very clean | Occupies a frequency range almost nothing else lives in |
| Drums | Clean | Transient signature is distinctive even under dense production |
| Vocals | Good to very good | Overlaps everything, but the model has seen the most vocal training data |
| Piano and guitar | Variable | Needs a six-stem model, and they mask each other in the midrange |
| Reverb tails | Poor | Reverb belongs to the room, not the source, so it smears across stems |
Heavily limited or clipped masters separate worse. The model was trained on music with intact dynamics, so a squashed source is out of distribution.
Getting a usable result on the first pass
- Feed it the highest-quality file you have. A 128 kbps MP3 has already thrown away the detail the model needs.
- Use a six-stem model when you need guitar or piano separately, and a four-stem model when you do not.
- Expect artefacts in dense sections and accept them, or edit around them.
- Check the separated stem in isolation and in context. Artefacts that are obvious solo often disappear in a mix.
- If the result is for a remix, work with it rather than against it. Filtering the stem is usually faster than fighting the model.
Where it is genuinely the right tool
- Remixing when you were never given the multitrack.
- Making an instrumental or an acapella from a finished master.
- Rescuing a part from an old bounce where the session is long gone.
- Practice and transcription, where you want to hear one part clearly.
- Karaoke and live backing tracks.
Where it is not the right tool: replacing a proper mix. If you have the multitrack, use the multitrack. Separation is a repair for a situation you did not choose.
A ten-minute version of this
If you have one evening and not one week, this is the order that gets the most improvement for the least time.
- Two minutes: listen all the way through without touching anything, and write down the three things that bother you.
- Two minutes: fix the loudest problem on the list, with the simplest tool that will do it.
- Two minutes: check in mono and on a phone speaker. Fix anything that falls apart.
- Two minutes: level-match against a reference and note the tonal difference, not the loudness difference.
- Two minutes: make one broad corrective move based on that comparison, then stop.
Three specific problems solved beats twenty small adjustments that cancel out. The written list is what stops the session drifting.
Knowing when to stop
There is a point in every session where further work stops improving the record and starts merely changing it. Recognising that point is most of the skill.
- You are undoing changes as often as you are making them.
- You cannot articulate what the last move fixed.
- The reference no longer sounds obviously better in any specific way.
- You are adjusting things you would not notice as a listener.
Bounce it, sleep on it, and listen once in the morning on a system you did not use to make it. That single listen tells you more than another two hours ever will.
Frequently asked questions
How accurate is AI stem separation?
Very good on bass and drums, good on vocals, and variable on piano and guitar. Quality depends heavily on the source: a dynamic, lossless master separates far better than a squashed MP3.
Can I get the original multitrack back?
No. The model estimates each source from the stereo mix. It is a reconstruction, not a recovery, and there is no hidden multitrack inside the file.
Why does the separated vocal sound watery?
That is a spectral artefact, and it shows up most in dense sections where the model has to guess. Feeding it a lossless, unlimited source reduces it noticeably.
Is it legal to use separated stems?
Separating audio you own or have licensed is fine. Releasing a remix built from someone else's record needs their permission, exactly as it always did.
Hear it on your own track
Upload a mix and get a mastered version back in minutes. Free to try, no card required.
Related tools: Stem Separation · Vocal Remover · Karaoke Maker
Related articles
What Are Stems and Why Mixing Needs Them? A Producer's Plain-English Guide
A practical guide to stems and why mixing needs them. What the model can actually isolate, why lossy sources hurt it,…
Vocal Remover vs Stem Separation: Which One Should You Use?
Vocal remover vs stem separation in practice: which sources come out clean, why a limited master separates worse, and…
Stem Separation Mixing Instrumental Karaoke: The Working Method
A practical guide to stem separation mixing instrumental karaoke. What the model can actually isolate, why lossy sourc…
Acapella Extractor Online: What Actually Matters
A practical guide to acapella extractor online. What the model can actually isolate, why lossy sources hurt it, and ho…
