Ten years ago, if you wanted the drum track from a finished song, the honest answer was that you could not have it. That recording sat in a studio archive, if it still existed at all. Today you can pull a usable approximation out of an ordinary MP3 in about a minute. This is a guide to what that actually is, what it is not, and the point at which it stops working.
First, what is a stem?
A stem is one layer of a song, kept separate from the others.
When a record is made, nothing is captured all at once. The drummer is recorded. The bass is recorded. Vocals are recorded, usually many times over, with the best moments stitched together afterwards. Each of those layers exists as its own file — the drum stem, the bass stem, the vocal stem.
At the end, a mixing engineer balances all of them and folds the result down into two channels: left and right. That two-channel version is what you stream or download. It is a photograph of a mix rather than the mix itself, and here is the part that matters — the folding is not reversible in any exact sense. Once two sounds have been added into a single number, no arithmetic can say with certainty what the two originals were.
Which is why separation is genuinely difficult, and why it sounded so poor for so long.
The problem, stated honestly
Imagine mixing yellow and blue paint. You get green. Now hand somebody that green paint and ask them to return the exact original yellow and the exact original blue. They can make a very educated guess. They cannot be certain, because endless combinations produce the same green.
Audio is the same problem, except with dozens of colours and tens of thousands of samples every second.
So separation never truly undoes a mix. It estimates — an informed reconstruction of what the layers most likely were. That single idea explains every strength and every limitation that follows.
How the old approach worked, and why it disappointed
The first widely available separation trick leaned on stereo geometry rather than any understanding of sound.
In a typical mix the lead vocal sits centred, appearing almost identically in both channels, while guitars and keys are spread wider. Subtract one channel from the other and anything perfectly centred cancels itself out. The voice largely vanishes.
Cheap, instant, and badly flawed. The kick drum, snare and bass are usually centred too, so they leave alongside the singer. The stereo image collapses into something narrow and boxy. Reverb tails spread across the stereo field survive as a ghost of the voice you were trying to remove. And on a mono recording there is no difference between the channels at all, so the method has nothing whatsoever to work with.
Most limiting of all, it can only ever produce two things: roughly-centred, and roughly-not-centred. It cannot hand you drums as their own track, because "drums" is not a position in the stereo field. It is a kind of sound.
What changed: recognition instead of geometry
Modern separation reframes the question entirely. Instead of asking what sits in the middle, it asks what a snare drum actually sounds like.
An engine is trained on a very large collection of music where the true separate layers were available. It sees the finished mix alongside the real answer, over and over. Gradually it learns the character of each instrument family — the sharp attack and fast decay of a snare, the way a bass note sustains underneath everything, the particular way a human voice moves between vowels and pauses to breathe.
Once an engine knows what things sound like, position stops mattering. It keeps the kick drum while removing the singer, even though both sit dead centre. It handles mono recordings. It finds a voice buried under distorted guitars occupying the same frequency range, because it is not judging frequency alone — it is recognising a pattern unfolding over time.
This is what the 7By.in Stem Splitter runs on, and it is why it returns four genuinely distinct stems rather than a crude centre-versus-sides split.
The stems you get
Separation typically returns four layers, and it is worth knowing what each actually contains.
Vocals. The lead voice, plus most backing vocals and harmonies — because those are voices too. If you were hoping harmonies would stay behind in the instrumental, they will not.
Drums. Kick, snare, hats, cymbals, percussion. Usually the cleanest of the four, because drums are sharp and distinctive, which makes them comparatively easy to identify.
Bass. The low-end instrument carrying the root notes. Generally reliable, though it can blur into a very low synth or a heavily sustained kick drum.
Other. Everything remaining — guitars, keys, strings, synths, brass. This is a leftover bin rather than a real category, and it is where most artefacts collect, simply because it absorbs whatever the other three did not claim.
You also get a full instrumental, which is just the mix with the vocal stem taken out. For karaoke that is usually the only file you need, and the Vocal Remover is the quicker route to it.
What separation is genuinely good at
On a cleanly produced modern record the results can be startling. Isolated drums are often usable in a production without anyone guessing their origin. Instrumentals pass as officially released versions. Isolated vocals come out clean enough to remix.
It performs best on material with space in it — where instruments are not fighting over the same frequencies, where the production was deliberate, and where the recording is reasonably recent.
Where it struggles
Being honest about this matters more than any marketing claim.
Dense, loud masters. Aggressively compressed music gives the engine less room to tell layers apart, because everything has been pushed to a similar level.
Live recordings. Crowd noise, stage bleed and room reflections blur exactly the boundaries the engine depends on.
Unfamiliar instruments. An engine trained mostly on drums, bass, voice and guitar has no confident slot for a sitar, a tabla or a regional folk instrument it has rarely encountered. Those land in "other", sometimes smeared.
Heavily processed vocals. A vocoder or extreme pitch manipulation makes a voice genuinely ambiguous. If a human listener cannot tell where the voice ends, an engine will not do better.
Poor source files. A low-bitrate MP3 discarded detail permanently when it was encoded. Separation cannot recover what is no longer there, and the difference against a WAV or FLAC is immediately audible.
What people use stems for
Producers and remixers get an acapella or a drum loop from a track that never had an official release of either. Musicians learning a part mute everything except the instrument they are studying — hearing a bass line completely alone, without a mix wrapped around it, is a genuinely different experience from trying to pick it out by ear.
DJs build mashups and smoother transitions. Teachers isolate a single line for a student to play along with. Content creators pull instrumentals for backing beds. Karaoke remains the most common use of all, and for that the instrumental is really all you need.
Audio restorers use it in a way most people never consider: separating a voice from background noise in an old recording, cleaning that one layer, then putting the pieces back together.
Practical advice
Start from the highest-quality file you own — this makes more difference than anything else within your control. Trim to the section you actually need first, using the Audio Cutter; shorter audio processes faster and costs fewer credits, since cost scales with length. Listen on headphones rather than laptop speakers, which hide precisely the low-frequency problems you are checking for. And resist re-separating an already-separated stem in the hope of cleaning it further, because artefacts compound rather than cancel out.
Common questions
Is this the same as vocal removal?
Vocal removal is one specific case of separation — take out the voice, keep everything else. Stem separation is the general version that also splits drums and bass into their own tracks.
Will I get the actual original studio stems?
No, and no tool can. Those exist only in the studio's archive. What you get is a reconstruction — frequently a convincing one, but an estimate rather than the master.
How many stems do I get?
Four separate layers, plus a combined instrumental.
Does it work on mono recordings?
Yes. Recognition-based separation does not rely on stereo difference, which is exactly why it succeeds where the old subtraction trick failed outright.
What does it cost?
You get 20 free credits every day. Separation costs 10 credits per five minutes, charged only when you download — previewing is always free.
Hear a song come apart
Four stems plus a full instrumental. 20 free credits every day.
Open the Stem Splitter →