Mixing and Mastering AI
Stem Separation

Stem Separation for Remix Workflow: The Working Method

Stem separation for remix workflow: how the model estimates each source, what separates cleanly and what does not, and how to get a usable stem on the first pass.

Mixing and Mastering AI · ·4 min read ·Stem Separation

This is one of those topics where the received wisdom is half right, which is worse than being wrong.

AI separation has moved from a novelty to a tool people ship records with. It is worth understanding what it estimates, because that tells you exactly where it will let you down.

Nothing below is theory for its own sake. Every number here is one you can check on a meter, and every step is one you can run on the next track you open.

Key takeaways

  • Separation estimates sources, it does not extract hidden multitracks.
  • Bass and drums separate most cleanly, piano and guitar least.
  • Feed it lossless audio, because lossy sources lose the detail the model needs.
  • Heavily limited masters separate worse than dynamic ones.

How AI stem separation actually works

A separation model is trained on many thousands of songs where the individual parts were already known. It learns what a vocal, a snare, a bass line and a piano look like in the time-frequency domain, then estimates a mask for each source in a mixed signal it has never heard.

It is estimation, not extraction. The original multitrack is not hidden inside the stereo file waiting to be recovered. The model is making an informed reconstruction, and that is why results vary with the material.

For vocals: The hardest source to isolate cleanly because the human voice overlaps almost every other instrument in the 200 Hz to 5 kHz range.

Stereo mixone fileModelmask estimateVocals92%typicalDrums96%typicalBass97%typicalOther78%typicalIt estimates each source. There is no hidden multitrack inside the file.
How a separation model turns one stereo file into isolated sources.

What separates cleanly, and what does not

SourceTypical resultWhy
BassVery cleanOccupies a frequency range almost nothing else lives in
DrumsCleanTransient signature is distinctive even under dense production
VocalsGood to very goodOverlaps everything, but the model has seen the most vocal training data
Piano and guitarVariableNeeds a six-stem model, and they mask each other in the midrange
Reverb tailsPoorReverb belongs to the room, not the source, so it smears across stems

Heavily limited or clipped masters separate worse. The model was trained on music with intact dynamics, so a squashed source is out of distribution.

Getting a usable result on the first pass

  1. Feed it the highest-quality file you have. A 128 kbps MP3 has already thrown away the detail the model needs.
  2. Use a six-stem model when you need guitar or piano separately, and a four-stem model when you do not.
  3. Expect artefacts in dense sections and accept them, or edit around them.
  4. Check the separated stem in isolation and in context. Artefacts that are obvious solo often disappear in a mix.
  5. If the result is for a remix, work with it rather than against it. Filtering the stem is usually faster than fighting the model.

Where it is genuinely the right tool

  • Remixing when you were never given the multitrack.
  • Making an instrumental or an acapella from a finished master.
  • Rescuing a part from an old bounce where the session is long gone.
  • Practice and transcription, where you want to hear one part clearly.
  • Karaoke and live backing tracks.

Where it is not the right tool: replacing a proper mix. If you have the multitrack, use the multitrack. Separation is a repair for a situation you did not choose.

Knowing when to stop

There is a point in every session where further work stops improving the record and starts merely changing it. Recognising that point is most of the skill.

  • You are undoing changes as often as you are making them.
  • You cannot articulate what the last move fixed.
  • The reference no longer sounds obviously better in any specific way.
  • You are adjusting things you would not notice as a listener.

Bounce it, sleep on it, and listen once in the morning on a system you did not use to make it. That single listen tells you more than another two hours ever will.

How it translates on real systems

Almost nobody hears your record on the system you made it on. The point of every decision above is that it survives the journey.

SystemWhat it exposesWhat to check
Phone speakerNo low end at all below roughly 500 HzDoes the bass line still read from its harmonics?
EarbudsExaggerated width and sub, hyped topIs anything sibilant or fatiguing?
LaptopThin midrange-only playbackIs the vocal still intelligible?
CarBoomy low end and road noiseDoes the low end turn to mush at 80 to 120 Hz?
Club systemEverything below 40 Hz, loudIs the sub mono and controlled?

If it works on a phone and in a car, it works almost everywhere else. Those two are the honest tests.

Frequently asked questions

How accurate is AI stem separation?

Very good on bass and drums, good on vocals, and variable on piano and guitar. Quality depends heavily on the source: a dynamic, lossless master separates far better than a squashed MP3.

Can I get the original multitrack back?

No. The model estimates each source from the stereo mix. It is a reconstruction, not a recovery, and there is no hidden multitrack inside the file.

Why does the separated vocal sound watery?

That is a spectral artefact, and it shows up most in dense sections where the model has to guess. Feeding it a lossless, unlimited source reduces it noticeably.

Is it legal to use separated stems?

Separating audio you own or have licensed is fine. Releasing a remix built from someone else's record needs their permission, exactly as it always did.

Hear it on your own track

Upload a mix and get a mastered version back in minutes. Free to try, no card required.

Related tools: Stem Separation  ·  Vocal Remover  ·  Karaoke Maker

Related articles