Glowing stereo waveform splitting into colour-coded stem layers that fan out into a 3D sphere of sound around a listener

From Stems to Immersive Audio: Turn Splits from Moises, LALAL.AI, AudioShake and Every Stem Splitter into Eclipsa Audio, ADM and Binaural

Spatial9 Team · · 19 min read

Key takeaways

  • Any stem splitter works: export your stems as same-length 24-bit WAV files from Moises, Music.ai, LALAL.AI, AudioShake, UVR, Demucs, your DAW or your DJ software, and Spatial9 places each one in 3D.
  • One upload gives you every immersive format: Eclipsa Audio (IAMF) for YouTube, ADM BWF for Atmos-style workflows, binaural for headphones, Ambisonics for 360 video, plus 5.1 and 7.1.4.
  • Open separation models improved from about 5 dB SDR (Open-Unmix, Spleeter, 2019) to about 9 to 12 dB (HT Demucs, BS-RoFormer, 2023), and Meta's SAM Audio (2025) can now pull out any sound you describe.
  • Separated stems carry small traces of each other (bleed), so keep core elements anchored front and centre and move the cleanest stems furthest; that one rule makes split stems sound great in 3D.

Short answer: if you already split songs into stems with Moises, Music.ai, LALAL.AI, AudioShake, Ultimate Vocal Remover, Demucs or any other stem splitter, you are one step away from immersive audio. Export the stems as WAV files of the same length, upload them to Spatial9, and our AI places every stem in a 3D room around the listener. From that one mix you can download Eclipsa Audio (IAMF) for YouTube, an ADM BWF master, a binaural headphone version, Ambisonics for 360 video, and 5.1 or 7.1.4 speaker mixes.

That is the whole workflow. The rest of this guide explains why it works, which stem splitters exist today, how the machine learning models behind them evolved over thirty years, and how to get the best possible 3D result from separated stems.

Workflow from any stem splitter through Spatial9 to Eclipsa Audio, ADM BWF, binaural, Ambisonics and 5.1 or 7.1.4
Any splitter in, every immersive format out. Stems become objects, objects become every format.

Why stems are the perfect starting point for immersive audio

A stereo song is two channels with everything glued together. An immersive mix is the opposite: each instrument is an independent audio object with its own position in 3D space, including above the listener. To place a guitar to your left and a pad above your head, the guitar and the pad have to exist as separate sounds first.

That is exactly what a stem splitter gives you. Each stem (vocals, drums, bass, guitar, keys, "other") becomes one object. Spatial9 then decides where each object belongs, how wide it should be, whether it should move, and how high it should sit, using the same placement rules a professional immersive mixer would apply.

If you have never mixed in 3D, our introduction to spatial audio covers the basics, and how to create spatial audio from stereo explains the full stereo-to-3D process.

The complete index of stem splitters in 2026

People search for "stem splitter" in many different ways, so here is a neutral index of the tools in use today, grouped by where you meet them. This is not a ranking: every tool on this list can produce stems that work with Spatial9. Most commercial tools do not publish which model they run, so we only discuss the underlying technology for the open-source models in the next sections.

Online apps and mobile apps

ToolWhere you use itTypical stems
MoisesWeb, iOS, Android, desktopVocals, drums, bass, guitar, piano, strings, wind, other
Music.aiWeb and developer platformConfigurable modules for vocals, instruments and dialogue
LALAL.AIWeb, desktop, mobileVocals, instrumental, drums, bass, guitars, piano, synth, strings, wind
AudioShakeWeb (AudioShake Indie) and developer platformVocals, instruments, dialogue, music and effects
FadrWeb and desktopVocals, drums, bass, melodies, plus MIDI
Gaudio StudioWebVocals, drums, bass, piano, guitar, other
Splitter.aiWebVocals, accompaniment, drums, bass
vocalremover.orgWebVocals and instrumental
PhonicMindWebVocals, drums, bass, other
Mobile karaoke and practice appsiOS and AndroidUsually vocals and instrumental

Inside DAWs and audio editors

ToolWhere you use it
Logic Pro Stem SplitterApple Logic Pro for Mac and iPad
FL Studio Stem SeparationImage-Line FL Studio
Ableton Live stem separationAbleton Live 12.3 and later
Steinberg SpectraLayersStandalone editor and Cubase integration
iZotope RX Music RebalanceiZotope RX repair suite
RipX DAW (Hit'n'Mix)Standalone stem-based editor
Acon Digital RemixPlug-in
AudionamixProfessional separation and dialogue tools
Studio One and other DAWsBuilt-in or plug-in based separation

Inside DJ software

ToolWhere you use it
Serato StemsSerato DJ Pro
rekordbox StemsPioneer DJ / AlphaTheta rekordbox
Traktor Pro stemsNative Instruments Traktor Pro 4
VirtualDJ StemsVirtualDJ
djay Neural MixAlgoriddim djay
Engine DJ StemsDenon DJ and Numark hardware

Free and open source

ToolWhat it is
Ultimate Vocal Remover (UVR)Free desktop app that bundles many open models (Demucs, MDX-Net, RoFormer and more)
DemucsMeta's open model family, runs from the command line or inside many apps
SpleeterDeezer's open model, the first widely used one
Open-UnmixInria and Sony open reference model
python-audio-separator and MSSTCommunity toolkits that run RoFormer, SCNet and MDX models
StemRollerFree desktop front end for Demucs
SAM AudioMeta's prompt-based separation model released in December 2025

A short history of music source separation

Separating a mix back into its instruments is one of the oldest problems in audio research. Engineers call it source separation, and the musical version is often nicknamed the "cocktail party problem for music". It took three decades to go from lab demos to one-click apps.

Timeline of music source separation from ICA in 1994 to SAM Audio in 2025
Three eras: classic signal processing, deep learning, and today's band-split transformers and prompted models.

1990s to 2016: the signal processing era

The first methods used pure mathematics rather than learning. Independent Component Analysis (ICA) in the mid-1990s assumed every source is statistically independent, which works for microphones in a room but not well for a finished stereo song. Non-negative Matrix Factorisation (NMF), popular from the early 2000s, broke a spectrogram (a picture of frequencies over time) into repeating building blocks that could be grouped into instruments. REPET (2012) exploited the fact that accompaniment repeats while vocals change. These methods were clever, but results were metallic and highly song-dependent.

2017: the benchmark that changed everything

Progress needs a fair test. In 2017 researchers released MUSDB18: 150 full-length songs (100 for training, 50 for testing) with professionally separated vocals, drums, bass and "other". A lossless version, MUSDB18-HQ, followed in 2019. Almost every model since then reports its quality on this dataset, and it is the reason the four-stem layout (vocals, drums, bass, other) became the industry default.

2018 and 2019: deep learning goes open source

Wave-U-Net (2018) showed that a neural network could separate audio working directly on the raw waveform. Then 2019 brought three open releases that put separation into everyone's hands:

2021: the Music Demixing Challenge

Sony and AIcrowd ran the Music Demixing (MDX) Challenge 2021, which pushed quality forward fast. Two important open models came out of it: Hybrid Demucs (Demucs v3) from Meta, which combined waveform and spectrogram processing, and KUIELab-MDX-Net from Korea University, which inspired a large family of community-trained "MDX" models distributed through Ultimate Vocal Remover.

2022 and 2023: transformers and band-splitting

HT Demucs (Hybrid Transformer Demucs, Demucs v4, 2022) added a transformer in the middle of Meta's hybrid design, letting the model look across many seconds of music at once. It reached about 9.0 dB SDR, and about 9.2 dB with sparse attention and fine-tuning. It is still the default model in many apps today.

At the same time, researchers rethought how to feed music to a model. Band-Split RNN (2022) cut the spectrum into bands that follow how instruments actually occupy frequencies. ByteDance combined that idea with a transformer using rotary position embeddings to create BS-RoFormer (2023), which won the Sound Demixing Challenge 2023 music track. It reported 9.8 dB on MUSDB18-HQ using the standard training set, and 12.0 dB with 500 extra training songs. Its sibling Mel-RoFormer (2023) uses bands spaced like human hearing (the mel scale) and is especially strong on vocals.

2024: lighter and faster

SCNet (Sparse Compression Network, 2024) reached about 9.0 dB with a fraction of the computation, showing that quality no longer needs huge models. The open community (UVR, MSST and others) fine-tuned RoFormer and SCNet checkpoints for specific targets such as lead versus backing vocals, de-reverb and individual instruments.

2025: prompted separation with Meta SAM Audio

In December 2025 Meta released SAM Audio, extending its "Segment Anything" family to sound. Instead of a fixed list of stems, you tell it what to extract: with a text description ("the electric guitar"), by clicking on the object in a video, or by marking a time span where the sound occurs. It returns the target sound and the remaining mix (the residual). Under the hood it is a generative diffusion transformer trained with flow matching, and Meta published the model so anyone can download and run it. Earlier language-queried research such as AudioSep (2023) pointed in the same direction, but SAM Audio made it practical for general audio.

How good are the models? Understanding SDR

Separation quality is usually measured with SDR (Signal-to-Distortion Ratio), in decibels. It compares a separated stem to the real studio stem: the higher the number, the closer they are. Every +3 dB is roughly a halving of the error energy, so the jump from about 5 dB in 2019 to about 10 dB in 2023 means the leftover error shrank to roughly a third.

Bar chart of average SDR on MUSDB18 for Open-Unmix, Spleeter, Demucs, Hybrid Demucs, Band-Split RNN, HT Demucs, SCNet and BS-RoFormer
Average SDR across vocals, drums, bass and other, as reported in each paper. Higher is better.
Open modelYearMade byApproachReported average SDRLicense
Open-Unmix2019Inria and SonySpectrogram mask, recurrent networkabout 5.3 dBMIT
Spleeter2019DeezerSpectrogram mask, U-Netabout 5.9 dBMIT
Demucs v22019 to 2020Meta AIWaveform U-Netabout 6.3 dBMIT
Hybrid Demucs (v3)2021Meta AIWaveform plus spectrogramabout 7.7 dBMIT
KUIELab-MDX-Net2021Korea UniversityTwo-stream spectrogram networkabout 7.5 dBMIT
Band-Split RNN2022Tencent AI LabBand-split plus recurrentabout 8.2 dBPaper and community code
HT Demucs (v4)2022Meta AIHybrid plus transformerabout 9.0 to 9.2 dBMIT
BS-RoFormer2023ByteDanceBand-split transformerabout 9.8 dB (12.0 dB with extra data)Paper; open community implementations
Mel-RoFormer2023ByteDanceMel-band transformerstrongest on vocals in its paperPaper; open community implementations
SCNet2024Academic researchSparse compression networkabout 9.0 dBMIT
SAM Audio2025Meta AIPrompted diffusion transformernot reported on fixed 4-stem SDRMeta's model license

The five families of separation models

Every model above belongs to one of five architectural families. Knowing the family tells you what a model is good at and what kind of artifacts to listen for.

Five families of separation architecture: spectrogram masks, waveform, hybrid, band-split and prompted
Each family fixed a weakness of the one before it.

1. Spectrogram mask models (Spleeter, Open-Unmix, MDX-Net)

These models look at the song as a spectrogram and predict a "mask" for each stem: which parts of the picture belong to the vocal, which to the drums. They are fast and light. Their weakness is phase: they usually reuse the phase of the original mix, so stems can sound slightly hollow or "underwater", especially on cymbals and reverb tails.

2. Waveform models (Wave-U-Net, Demucs v1 and v2)

These skip the spectrogram and work on raw samples. They keep transients (the sharp start of a drum hit or a plucked note) crisp, which is why early Demucs was loved for drums and bass. Their weakness is high-frequency detail and occasional noisy artifacts.

3. Hybrid models (Hybrid Demucs, HT Demucs)

Meta's answer was to run both views side by side, waveform and spectrogram, and let them share information. HT Demucs adds a transformer between the two branches so the model can use long musical context: if the chorus repeats, it uses the first chorus to separate the second one better. The result is a balanced all-rounder, which is why it remains the default in so many apps.

4. Band-split models (Band-Split RNN, BS-RoFormer, Mel-RoFormer, SCNet)

Instruments live in different frequency ranges, so these models cut the spectrum into bands first, then model each band over time and the bands against each other. With a transformer on top (the "RoFormer" part) they currently give the cleanest vocals and the least bleed. Their trade-off is compute: the largest RoFormer checkpoints are slow without a good graphics card.

5. Prompted models (AudioSep, SAM Audio)

Instead of a fixed list of four or six stems, you ask for a sound. This is a big shift for immersive audio: you could pull out "the tambourine", "the crowd noise" or "the string section" as separate objects. Because these models are generative, they can sometimes invent small details rather than only subtracting, so they are best used for specific targets alongside a classic stem split.

What each model can actually pull out

The number and type of stems matters more for 3D mixing than a fraction of a decibel. Every extra stem is another object you can place in space.

Matrix of which stems Spleeter, Open-Unmix, Demucs, MDX-Net, BS-RoFormer, SCNet and SAM Audio can output
Four stems is the classic baseline; prompted models remove the limit entirely.

The nuances: what separation does to your sound

Separated stems are estimates, not studio originals. Understanding the typical side effects helps you get a great 3D result.

Bleed

A little of each instrument always remains in the other stems: a ghost of the hi-hat in the vocal, a hint of vocal in the guitars. In stereo you never notice, because all stems play from the same two speakers. In 3D the stems are pulled apart, and the ghost can suddenly appear from the "wrong" direction.

Illustrative chart of how audible stem bleed becomes as sources are placed further apart, for older splits, modern splits and original studio stems
The cleaner the stems, the further apart you can place them before ghosts appear. Illustrative.

This is the most important nuance of all, and it is why Spatial9 analyses the stems before placing them: stems with more bleed are kept closer to their partners, while clean stems are allowed to travel further.

Phase and "underwater" sound

Mask-based models reuse the phase of the original mix, which can make individual stems sound a little hollow when soloed. When the stems play together again the effect mostly cancels out, which is why we recommend keeping all stems in the mix rather than muting some.

Transients

Drum hits and plucked strings are short and dense. Older spectrogram models smear them; waveform and hybrid models keep them sharp. If your drums sound soft after splitting, try a Demucs or RoFormer based model.

Reverb tails

Reverb belongs to every instrument at once, so models split it unpredictably: some of it lands in "other", some stays with the vocal. In an immersive mix this is actually useful: the reverb that ends up in "other" can be lifted into the height speakers to create a sense of room.

The "other" stem

"Other" is a catch-all for everything that is not vocals, drums or bass: guitars, keys, synths, strings, effects. It is usually the widest and most complex stem, so Spatial9 treats it as a broad, enveloping object rather than a single point.

Step by step: from any stem splitter to immersive audio with Spatial9

Step 1: Split your song

Use whichever tool you prefer. If you can choose the model, a recent band-split or hybrid model (RoFormer or HT Demucs) gives the cleanest result for 3D. Choose the highest number of stems available.

Step 2: Export correctly

Six-point checklist for preparing stems: same start point, same length, WAV 24-bit 48 kHz, no master bus processing, clear names, check the sum
Two minutes of preparation saves hours of fixing timing problems later.

Our guide to exporting audio stems goes deeper on each point.

Step 3: Upload to Spatial9

Open Spatial9 and drop in your stems. Each stem is analysed: instrument, energy, frequency range, how much bleed it carries, and its role in the song. No stems? Upload the stereo master and Spatial9 separates it first. You can read more on the stems-to-spatial solution page.

Step 4: Let the AI place every stem

GravityLLM, our placement engine, builds a 3D scene the way an experienced immersive mixer would: core elements anchored at the front, harmony instruments spread to the sides, air and ambience in the height layer, low frequencies kept centred. You can see and hear every object in the 3D view and move anything you want.

Top-down room map showing where separated stems usually land in an immersive mix: vocal and kick centre, guitars and keys at the sides, pads, cymbals and reverb raised
A starting map for separated stems. Dashed rings show elements raised into the height layer.

Step 5: Export every format you need

FormatBest forWhere it plays
Eclipsa Audio (IAMF)YouTube immersive audioYouTube on supported TVs, browsers and devices, with binaural fallback on headphones
ADM BWFAtmos-style mastering, distribution and archivingProfessional immersive workflows, renderers and DAWs
BinauralHeadphones and earbudsAny player, any streaming service, any pair of headphones
Ambisonics360 video, VR, AR and gamesYouTube 360, VR headsets, game engines
5.1Film, TV and home cinemaAV receivers, soundbars, broadcast
7.1.4Immersive speaker roomsHome theatres and studios with height speakers

Which stems work best for each format?

Tips for the best result from split stems

  1. Split once, from the best source. Start from a lossless master, not a streaming rip or MP3.
  2. Prefer more stems over a cleaner but smaller set, as long as each stem still sounds musical.
  3. Keep vocal, kick, snare and bass anchored. They carry the song's punch and should feel solid on every system.
  4. Lift the ambience. Reverb, pads, backing vocals and cymbals sound most immersive slightly above the listener.
  5. Listen on headphones and speakers. Bleed that is invisible on speakers can become audible on headphones.
  6. Respect the rights. Separating a song does not change who owns it; only publish immersive versions of music you own or are licensed to use.

Why this matters now

For years immersive audio required original multitrack sessions, an expensive studio and a specialist engineer. Two things changed. Open separation models crossed the quality line where stems are clean enough to move around a 3D room. And open immersive formats such as Eclipsa Audio arrived on YouTube, the largest music platform in the world. Spatial9 connects the two: any stem from any splitter becomes an immersive mix for every format, in minutes. Our partnership with AudioShake, announced at CES 2025, was an early step on this path.

Frequently asked questions

Can I use stems from Moises, LALAL.AI, AudioShake or Music.ai with Spatial9?

Yes. Spatial9 accepts stems from any stem splitter, including Moises, Music.ai, LALAL.AI, AudioShake, Ultimate Vocal Remover, Demucs, Logic Pro, FL Studio, Ableton Live, Serato, rekordbox and others. Export them as WAV files with the same start point and length, then upload them together.

What is the best open-source model for stem separation?

On the standard MUSDB18 benchmark, band-split transformer models such as BS-RoFormer and Mel-RoFormer report the highest scores, followed closely by Meta's HT Demucs and SCNet. HT Demucs remains the most widely used all-rounder, while RoFormer models usually give the cleanest vocals.

What is Meta's SAM Audio?

SAM Audio is a prompt-based sound separation model released by Meta in December 2025. Instead of a fixed set of stems, you describe the sound you want with text, by clicking on an object in a video, or by marking a time span, and it extracts that sound plus the remaining mix.

How many stems do I need for immersive audio?

Four stems (vocals, drums, bass and other) are enough for a convincing immersive mix. Six or more stems, such as separate guitar, piano, backing vocals or reverb, give more objects to place and a more detailed 3D result.

Can I turn stems into Eclipsa Audio for YouTube?

Yes. Upload your stems to Spatial9 and export Eclipsa Audio (IAMF), the open immersive format supported by YouTube. Viewers with compatible devices hear the immersive mix, and headphone listeners get a binaural rendering.

Can I get an ADM BWF file from separated stems?

Yes. Spatial9 places every stem as an object and exports an ADM BWF master that can be opened in professional immersive workflows, alongside binaural, Ambisonics, 5.1 and 7.1.4 versions.

Do I need stems at all, or can I upload a stereo song?

You can upload a stereo song. Spatial9 separates it into stems automatically before placing them in 3D. If you already have stems from a splitter or from the original session, uploading them gives you more control.

Does stem bleed ruin an immersive mix?

No, but it changes how far apart stems should be placed. Spatial9 detects bleed and keeps the affected stems closer to their partners, so ghost sounds do not appear from unexpected directions.

What does SDR mean in stem separation?

SDR, or Signal-to-Distortion Ratio, measures in decibels how closely a separated stem matches the real studio stem. Higher is better; every 3 dB roughly halves the remaining error. Open models went from about 5 dB in 2019 to about 10 dB in 2023.

Sources

  1. Rafii et al., MUSDB18: a corpus for music separation (2017)
  2. Stoller, Ewert and Dixon, Wave-U-Net (ISMIR 2018)
  3. Hennequin et al., Spleeter: a fast and efficient music source separation tool (JOSS 2020)
  4. Stöter et al., Open-Unmix: a reference implementation for music source separation (JOSS 2019)
  5. Défossez et al., Music Source Separation in the Waveform Domain (Demucs, 2019)
  6. Défossez, Hybrid Spectrogram and Waveform Source Separation (2021)
  7. Kim et al., KUIELab-MDX-Net (2021)
  8. Rouard, Massa and Défossez, Hybrid Transformers for Music Source Separation (HT Demucs, 2022)
  9. Luo and Yu, Music Source Separation with Band-Split RNN (2022)
  10. Lu et al., Music Source Separation with Band-Split RoPE Transformer (BS-RoFormer, 2023)
  11. Wang et al., Mel-Band RoFormer for Music Source Separation (2023)
  12. Fabbro et al., The Sound Demixing Challenge 2023, Music Demixing Track (2023)
  13. Tong et al., SCNet: Sparse Compression Network for Music Source Separation (2024)
  14. Liu et al., Separate Anything You Describe (AudioSep, 2023)
  15. Meta AI, SAM Audio (2025)
  16. Demucs repository
  17. Ultimate Vocal Remover repository

Try Spatial9 free