AI Spatial Audio Mixing vs Manual and Plugins: What Spatial9 Automates
Luiz Zanardo (CEO & Founder) ·
It is 9:00 on a Monday. A five-minute pop single lands in your inbox: 16 stems, 120 BPM, two verses, three choruses, a bridge, a breakdown and a big final lift. The label wants an immersive version by Friday.
You open the session. The sphere around the listener is empty. Every voice, every synth and every drum is waiting for you to tell it where to live, how to move, when to rise and when to fall back.
Three engineers get the same song that morning. The first does everything by hand. The second has the best spatial plugins on the market. The third uses Spatial9.
By Friday, only one of them has spent their week listening instead of drawing. Let me show you why, one task at a time.
One 5-minute pop song, three ways to mix it in 3D
Here is the song behind every number in this article: 5 minutes, 120 BPM in 4/4, 150 bars, 600 beats, 16 stems, about 400 sung words and roughly 300 chord events. Ten stems move. The kick, the bass and the lead vocal stay anchored in their main position.
The three scenarios:
- A. Fully by hand. A skilled immersive engineer places, moves and checks every stem with standard panners and automation lanes.
- B. Engineer with today's plugins. The same engineer uses the best current tools: procedural motion generators, tempo-synced and audio-reactive panners, prompt-based motion tools, upmixers and headphone rendering built into the music software.
- C. Spatial9. The song is analysed and choreographed automatically. The engineer uploads, listens, reviews and decides.
| Task | A. By hand | B. With plugins | C. Spatial9 (human time) |
|---|---|---|---|
| Session prep, stem import and gain staging | 1.5 h | 1.5 h | 10 min (upload and settings) |
| Base placement of 16 stems in 3D | 1.5 h | 1 h | Automatic |
| Movement design for 10 stems over 600 beats | 10 h | 3 h | Automatic, from 1,500+ movements |
| Section staging across about 10 section changes | 6 h | 4 h | Automatic, nine section profiles |
| Beat-locking and per-stem phase offsets | 3 h | 1 h | Automatic |
| Harmonic motion across about 300 chord events | 12 h | 10 h | Automatic |
| Lyric choreography: about 400 words, about 40 cues | 8 h | 7 h | 15 min to approve or reject |
| Perception-driven effects design | 8 h | 5 h | Pick from 61, heard in review |
| Frequency-aware placement and tuning | 4 h | 3 h | Automatic |
| Vocal clarity and ducking | 2 h | 1.5 h | Automatic |
| Headphone, head-tracking and stereo A/B checks | 2 h | 2 h | 45 min of listening |
| Two rounds of revisions | 6 h | 4 h | 30 min of notes and tweaks |
| Exports for six formats | 6 h | 3 h | 10 min to choose formats and spot-check |
| Total active human time | About 70 h | About 46 h | About 1 h 50 min |
Two notes on fairness. Listening checks of each exported format sit inside the 45-minute review block, and destination-specific delivery rules remain your responsibility in all three scenarios. The plugin column assumes a skilled engineer who already knows the tools well.
What today's plugins genuinely solve
Credit where it is due. Today's spatial tools are good, and scenario B is a real improvement:
- Procedural motion generators create paths for you instead of making you draw every point. Movement drops from 10 hours to about 3.
- Tempo-synced and audio-reactive panners move sound with the beat, loudness or brightness of a single track.
- Prompt-based motion tools turn a written idea into motion, and some can coordinate many objects at once.
- Upmixers spread stereo or surround material into an immersive layout quickly.
- Headphone rendering inside the music software lets you preview a binaural version without a speaker room.
What they still leave on your desk
Existing tools can generate complex motion, coordinate multiple objects, react to audio and automate parts of authoring and delivery. What they do not do in our model is connect those decisions to the song's structure, harmony, lyrics and the role of each stem in one workflow. So in scenario B, the engineer still has to:
- Follow about 300 chord events and decide how the field should respond to each one. That is why harmony only drops from 12 hours to 10.
- Read the lyrics and choose which words deserve a spatial gesture. That only drops from 8 hours to 7.
- Make ten section changes feel coordinated across 16 stems, not just move one stem well.
- Pick, tune and combine effects, and keep all of it consistent through two rounds of revisions.
Plugins speed up the drawing. They do not take over the decisions. Spatial9 does both, and leaves you the final word.
In real life, plenty of mixes skip the harmonic motion, the lyric choreography and the perception design, sometimes by choice and often for lack of time. Spatial9 makes that a creative decision instead of a budget decision.
The cost model: three scenarios side by side
To keep the comparison fair, every scenario uses the same rate: $100 per hour for a skilled engineer, plus $35 per hour for the room, monitoring, computer and software, so $135 per working hour. Spatial9 human time is charged at that same full rate. Scenario B also includes about $50 per track for specialised spatial plugins, assuming roughly $1,500 a year of plugin licences spread over 30 immersive songs. General music software is already inside the $35 hourly overhead, so it is not counted twice. Scenario C includes the Spatial9 price: $9.90 for the entire track, one time. Everything else in that column is the engineer's own listening and review time.
| Scale | A. By hand | B. With plugins | C. Spatial9 |
|---|---|---|---|
| One song | 70 h · $9,450 | 46 h · about $6,260 | About 1.8 h · about $257 ($9.90 Spatial9 + review time) |
| 10-song album | 700 h · $94,500 | 460 h · about $62,600 | About 18 h · about $2,570 ($99 Spatial9 + review time) |
| 100 songs a year | 7,000 h · $945,000 | 4,600 h · about $626,000 | About 183 h · about $25,700 ($990 Spatial9 + review time) |
This is a cost comparison, not a promise of profit. So here is what the numbers actually mean, in three separate ways:
- Capacity released. Spatial9 frees about 68 engineer hours per song compared with doing it by hand, and about 44 hours compared with a plugin-assisted workflow. Across 100 songs that is the workload of about 3.6 full-time engineers versus by hand, or about 2.3 versus plugins.
- Spend avoided. Under these assumptions the modelled cost is about 97% lower than by hand and about 96% lower than with plugins, roughly one thirty-seventh and one twenty-fourth. And Spatial9 itself costs $9.90 per track, less than five minutes of the engineer's time.
- Real financial return. Saved hours only become money when they are used for something else: more releases, more clients, or work you no longer have to buy in. At $135 per hour, the $9.90 track price is covered as soon as Spatial9 frees less than 5 minutes of time you would otherwise pay for or bill. If the hours sit idle, the saving is time, not cash.
The total cost of ownership also depends on what you need to deliver. A headphone-first creator can do everything on a laptop. A professional delivery still deserves a final check in a proper speaker room. Spatial9 changes how many hours you need that room for, not whether good monitoring matters.
Do you still need a DAW, plugins and a renderer?
Here is the part of the bill most people never add up. Before an engineer can spend the first of those 70 or 46 hours, someone has to buy the tools.
A traditional immersive workflow is a stack: music software (a DAW) with immersive support, a shelf of spatial plugins for panning, upmixing and effects, a renderer and encoder for each delivery format, a room with a dozen or more speakers, acoustic treatment, a powerful computer, and the training to make all of it work together.
To be fair, not everyone needs the whole stack. Some music software already renders immersive mixes on headphones, so a creator can start on a laptop. The honest comparison depends on who you are:
| Who you are | Traditional stack | With Spatial9 |
|---|---|---|
| Headphone-first creator | Music software with immersive support plus a few spatial plugins, about $500 to $3,000 | Good headphones, about $200 to $500, then $9.90 per track |
| Professional delivery | A sample full immersive room, about $18,000 to $86,000 in year one, or rented room time | Create and review on headphones, then book room time for the final check |
| Catalog or AI music platform | Engineers and rooms that scale with every track | The engine and API scale with the catalog; people review, not draw |
| Sample professional room budget | Typical year-one range | With Spatial9 |
|---|---|---|
| Music software with immersive support | $200 to $600 | Optional, for engineers who want to refine further |
| Spatial panners, upmixers and immersive plugins | $500 to $3,000 | Built in: 1,500+ movements and 61 effects |
| Renderer and encoder licences | $300 to $1,500 | Built in: rendering and up to 33 export targets |
| 7.1.4 speakers, subwoofers and multichannel interface | $8,000 to $40,000 | Not needed to create; good headphones, $200 to $500 |
| Acoustic treatment for an immersive room | $5,000 to $30,000 | Not needed to create |
| Local processing power for large object sessions | $3,000 to $6,000 | A device you already own and an internet connection; processing runs on our servers |
| Training and calibration | $1,000 to $5,000 | Plain language and a guided assistant |
| Total before any engineer hour | About $18,000 to $86,000 | About $200 to $500 for headphones, then $9.90 per track |
Then the stack keeps charging you. Subscriptions renew, plugins need paid upgrades, renderers change versions, computers age out, and every new delivery format means another tool to learn. That is the $35 per hour we built into the cost model for room, monitoring and software. Renting a room is an option on both sides; the real difference is how many hours of it you need.
With Spatial9, the creative work lives in one place. Movement, effects, harmony, lyrics, rendering, checks and export all happen in the same engine. You need ears, good headphones and taste. A final listen on speakers is still a good idea for a professional release.
The new age of AI music does not have a mixing desk
This matters even more for where music is going.
A growing share of new music is now created with AI music platforms. A song can be generated, rewritten and finished in minutes, often without anyone ever opening a session. These platforms produce enormous volumes of music, and their users expect results at the same speed.
There is no DAW in that workflow. No plugin chain, no renderer, no engineer booked three weeks out. The economics of the old way simply do not work: at 70 hours per song, a catalog of 10,000 AI-generated tracks would need 700,000 engineer hours, about 350 engineers working for a full year. Even with today's best plugins, at 46 hours per song, it is still 460,000 hours.
This is the role Spatial9 was built to play. We are the spatial layer of the AI music stack: a song or its stems goes in, through the app or our API, and a living, intentional immersive version comes out, ready to export. We have already done it at scale with Pozalabs, transforming more than 1,000 AI-generated tracks into immersive audio.
For AI music, immersive sound stops being a premium add-on reserved for a few releases. It can become a feature of every track.
What about the engineers who love their tools?
They are welcome, and they stay in control. Spatial9 exports ADM BWF, stems with metadata and the major immersive formats, so a professional can take the engine's work into their own session and keep refining it. The difference is that the tools become a choice, not an entry ticket.
And one rule does not change with any stack: every destination has its own delivery requirements. Some stores, Apple Music among them, require Dolby Atmos music to come from original multitracks or stems made from them. Check the rules for where your music is going, whichever way you make it.
Who is speaking, and exactly what I am claiming
I have spent my career around people who mix for a living, and I have deep respect for what they do. So let me say this carefully, and then let me say it boldly: the immersive mix that music deserves is enormously hard to draw by hand. Not because engineers lack talent, but because a 3D mix is not one decision. It is thousands of decisions, made in sync with the beat, the sections, the harmony and the words of a song, all at once.
Let me also be precise about what I am claiming. Spatial9 did not invent spatial automation. Upmixers, audio-reactive panners, procedural motion generators, prompt-driven motion tools and full immersive authoring suites already exist, and many of them are excellent. What we built is a different approach: one connected engine where the structure, the harmony, the lyrics and the way we perceive sound all shape a song's space together, and where every decision stays in the hands of the artist.
In this article I will walk through every major feature, with the real numbers inside it, and compare each one honestly with what it takes to do the same by hand or with plugins. Some of it is merely slow by hand. Some of it is so time-consuming that, in practice, almost nobody does it.
The honest truth about mixing in 3D by hand
In stereo, a mix lives on a line between two speakers. In immersive audio, every sound lives somewhere in a sphere around the listener, and it can move. That freedom is the whole magic, and it is also the trap.
Go back to our 5-minute song. Ten moving stems, each with a position in three dimensions, is 30 lanes of automation before you touch reverb, level or EQ. One point per beat across 600 beats is 18,000 automation values. Motion generators can draw many of those points for you, and that is real help. But the hard part was never the drawing. It is deciding how each move should agree with the beat, the section, the chord and the lyric being sung at that moment.
Some projects benefit from rich spatial choreography; others benefit from restraint, and a still arrangement can be a deliberate, beautiful choice. The problem is when stillness is forced by the clock rather than chosen. Spatial9 is designed to make either choice free of repetitive setup.
Here is how, feature by feature.
1. More than 1,500 movements, built on 110+ core shapes
Our movement library holds more than 1,500 movements. At its heart are more than 110 core movement shapes in 20 families, each with a described purpose and psychoacoustic effect. Around that core sit roughly 1,350 movements designed for specific instruments and traditions: plucked strings, winds, idiophones, drums, bowed strings, mallets, brass, free reeds and keyboards, from the Afghan rubab and Persian santur to the oud, the ocarina and the West African talking drum.
A taste of the core families:
- Circular orbits (11): from a slow orbit and a micro orbit to a precessing orbit and an orbit with stops.
- Spirals and corkscrews (11): spiral in, spiral out, throw, suck in, rising and dropping corkscrews.
- Musical gestures (8): Crescendo Approach, Diminuendo Retreat, Staccato Bounce, Legato Flow, Trill Flutter, Glissando Sweep, Arpeggio Climb and Chord Bloom.
- Psychoacoustic (5): Doppler Pass, Precedence Ping, Haas Spread, Binaural Beat and Shepard Tone Rise.
- Cinematic (7): Hero Entrance, Villain Circle, Jump Scare, Epic Flyover, Chase Sequence, Slow Reveal and Impact Shockwave.
- Nature (8): Butterfly Wings, Bird Flight, Bee Dance, Falling Leaf, Water Ripple, Ocean Wave, Wind Gust and Rising Smoke.
- Architectural (4): Cathedral Dome, Corridor Echo, Amphitheater Rise and Balcony Overlook.
Under the hood, each movement is a real mathematical path in three dimensions, not a canned pan. Paths respect a minimum distance of 0.8 metres from the listener, a soft boundary that keeps sound inside the room, and a vertical range scaled to the radius. Every stem gets its own phase, so two instruments sharing an orbit stay out of step instead of overlapping.
By hand: drawing a convincing figure-8, a Mobius strip or a falling leaf means plotting dozens of points per cycle on three lanes and keeping them in time. One stem is an afternoon. A library of more than 1,500 gestures, each tuned to the character of an instrument, is something no studio could ever build session by session.
2. Every section of the song gets its own space
A song is a story with chapters. Our engine reads the structure and gives each chapter its own size, height and kind of motion.
- Intro: compact and slightly below ear level (expansion 0.75, height -0.5 m). The room is still waking up.
- Verse: a little wider and just above ear level (0.85, +0.15 m). Intimate, close to the storyteller.
- Buildup: the field grows with the energy of the build itself, rising and widening bar by bar.
- Chorus: the room opens (1.4) and lifts 1.2 metres. This is where the listener feels surrounded.
- Bridge: a change of perspective (1.1, +0.6 m), wide enough to feel new without stealing the chorus.
- Drop: maximum impact (1.5, +1.5 m).
- Breakdown or break: the space pulls in (0.7) and sinks below ear level (-0.8 m), so the return hits harder.
- Climax: the widest and highest point of the song (1.6, +1.8 m).
- Outro: the room gently closes (0.8, -0.3 m) and lets the listener go.
Inside each section, moving stems choose patterns that match the energy: spirals, figure-8s, zigzags and bounces when it is high, and breathing, floating and swaying when it is low. Atmospheric stems stay soft and wave-like. Changes between sections ease in over one to four bars, depending on the profile you choose.
Equally important is what does not move. The kick, the bass and the lead vocal stay anchored, so the song keeps its spine while the space around it breathes.
By hand: building nine section-aware states and moving every stem between them musically is a full mixing project in itself. Redoing it every time the arrangement changes is where manual immersive mixing usually gives up.
3. Harmony moves the room, subtly and constantly
This is the feature people feel before they can name it. The engine estimates harmonic features, such as key, mode and chords, and maps them into adjustable spatial behaviours. These mappings are creative interpretations, not measurements of a listener's emotional response.
It analyses the spectrum between 60 Hz and 4,000 Hz, folds it into the 12 pitch classes and matches the result against major and minor key profiles to find the key. It then estimates each chord, with its quality, scale degree, tension and consonance. Each mode carries a brightness, from Locrian at 0.00, through Phrygian 0.17, Minor 0.33, Dorian 0.50, Mixolydian 0.67 and Major 0.83, to Lydian at 1.00.
Then come the subtle movements, the ones that make a mix feel alive rather than animated:
- Brightness shapes the room. Brighter modes make the field larger and raise its ceiling. A Lydian passage is mapped to a more open field than a Phrygian one.
- Harmonic breathing. The whole field breathes in a slow two-bar cycle, and each chord colours the breath. The dominant (V) leans outward with anticipation, the subdominant (IV) opens slightly, the relative minor (vi) draws a little inward, and the tonic (I) settles home.
- Key-change warps. When the song modulates, the space reacts over two seconds. A half-step lift expands the field and turns it by about 30 degrees, like a truck-driver key change you can see. A move to the relative major or minor gently contracts it. A tritone shift swirls. A move up a fifth turns the room just a little.
- A circle-of-fifths compass. Each key has its own orientation in the room, so harmonic journeys become journeys through space.
- Consonance and tension. Consonant, settled chords let sounds gather, while tension pushes them apart, with a settling time so nothing jitters.
- Scale-degree placement. The melody note's function nudges direction and distance by small amounts.
- Key character. Each key carries a character inspired by Baroque key theory, adding intensity, vertical lean and symmetry.
All of this is deliberately small, smoothed continuously and scaled by section energy. You do not hear things jumping around. You feel the harmony in the space.
By hand: a gifted engineer feels this intuitively, but tracking every chord, cadence and modulation and translating each one into gentle three-dimensional motion across a whole session would take days. Even with the best plugins this stays manual. It is the clearest case where the time cost decides the outcome, so in practice it rarely gets done. The values the engine uses are creative choices, not laws of music: they are tuned to sound natural, and you can change or switch them off.
4. The voice leads: lyrics that choreograph the space
This is one of my favourite parts of the engine, because it connects the space to the meaning of the song.
Here is the chain:
- The vocal stem is transcribed word by word, with a timestamp for every word.
- The lyrics are read for meaning, using 154 built-in words and phrases in six families (direction 39, emotion 28, compound phrases 27, nature 23, time 20 and distance 17) plus an AI reading of the context around them.
- Each match becomes one of 22 spatial actions, including elevation, pan, distance, scatter, collapse, orbit, freeze, isolation, spiral, wave, sweep, dome, breathe, pulse and anchor.
- The action lands exactly on the word.
When the singer says "falling", the sound descends. On "alone", the vocal is isolated. On "together", everything collapses inward. On "rising up", the space lifts overhead. You choose whether the cues shape the vocal layers only (doubles, harmonies, echoes and small gestures around a centred lead) or the whole mix, and you approve, edit or reject every single move.
We made one deliberate choice: the lead vocal itself stays anchored at the front for focus and intelligibility, while the singer's words lead what happens around it. The voice becomes the conductor, not a fly buzzing around your head. Vocal-aware ducking keeps room for the singer whenever they are present.
By hand: an engineer would need to transcribe the song, time every word, decide which words carry spatial meaning, and draw a move for each one on the exact syllable, across every affected stem. For one song that is days of work. Spatial9 makes lyric-timed spatial suggestions part of the workflow, with every move open to approval, editing or rejection.
5. 61 effects inspired by how we hear
Spatial9 combines conventional processing with experimental spatial effects inspired by psychoacoustics. Their names describe creative intentions; what you perceive depends on the source, the settings, the renderer, the playback system and the listener.
The library holds 61 effects in nine families. Six are the traditional tools you know, rebuilt to be spatially aware: reverb, EQ, delay, dynamics, width and distance. The other 55 are experimental effects inspired by perception research and acoustic science, including classic psychoacoustic discoveries such as the Shepard tone and the precedence and Haas effects.
- Psychoacoustic illusions (9): a Phantom Source Generator that places sound where no speaker exists, a Precedence Puppeteer that snaps objects to unexpected directions with micro-delays, Looming Gravity that makes sounds approach as musical tension rises, a Continuity Teleport Mask that moves objects across the room while your brain insists the motion was continuous, an Externalization Handshake that pulls sound out of your head on headphones, and Motion Afterimage Painting, where still sounds seem to drift after a fast swirl.
- Physics simulation (7): a Metamaterial Refraction Room that bends sound around the listener like a lens, a One-Way Sound Gate, a Sonic Mirage Horizon that floats distant sounds above you, turbulence, material propagation, air-density layers and gravity wells.
- Dimensional (8): a Mobius Panner whose orbit flips left and right, a Klein Bottle Room where sound leaks from inside to outside, a Tesseract Teleporter, fractal spatialization and kaleidoscope clones.
- Temporal (6): Beat-Snap Wormholes that jump on the downbeat, time-reversal reverb that retraces its path, retrograde echoes and temporal smearing.
- Reactive (11): Dynamic Masking Avoidance that moves sounds out of each other's way, Harmonic Attraction that pulls consonant sounds together, transient bursts, an energy-responsive room size and phrase-quantized movement.
- Biological, sensory and narrative (14): breath-synced and heartbeat fields, an infrasonic presence field, a memory palace where returning themes return to the same room, and an unreliable narrator where space deliberately contradicts harmony.
By hand: some of these illusions can be built manually by an expert with a lab-grade setup, one at a time. A few have plugin cousins, many would need custom setups, and combining them with movement, harmony and lyrics in one coordinated mix is not what a chain of separate plugins was designed to do.
6. Movement from a quick idea
You type what you imagine, in plain words, and the engine gives you a movement.
Write "a leaf falling slowly, then caught by a gust" and GravityLLM, our spatial intelligence, interprets the feeling and the shape you mean. It returns a real path with formulas for x, y and z over time, a name and one of six categories: circular, linear, organic, rhythmic, atmospheric or complex. Not quite right? Keep the conversation going and refine it.
Every generated movement is held within a speed of 0.5 to 2.0, a radius of 1.5 to 4 metres and a path within plus or minus 4 on every axis. Your idea gets bold, but it never gets broken.
By hand: turning a poetic image into a smooth three-dimensional curve means sketching it, translating it into keyframes and fixing the jerks. Other tools now offer prompt-driven motion too. Our difference is that the result drops straight into a session that already understands the song's sections, harmony and lyrics.
7. Perform the movement with your hands
Sometimes words are not enough and you want to show it. Spatial9 includes hand tracking and full-body pose tracking from your camera. Draw a shape in the air or move your body, and the gesture becomes a movement you can put on any stem.
By hand: ironically, this is the most human feature we have, and a mouse and automation lane cannot capture the arc of a gesture in three dimensions.
8. Beat-locked motion, with taste built in
Movement that ignores the groove sounds like a screensaver. In Spatial9, the cycle of an orbit can be set in bars, phases offset in beats, and section changes eased over one to four bars.
Transitions follow a three-phase curve: anticipation, crossover and settle, with the crossover complete at 75% of the move. Speed is capped and every move lasts at least half a second, which keeps motion smooth rather than whipping past your ears. Pads, atmospheres and effects get an extra slow drift of a few centimetres over 16 to 64 bars, the kind of life you feel more than notice.
By hand: keeping hundreds of moves locked to tempo, with consistent anticipation and settle, is the definition of tedious, and it is the first thing to fall apart when the tempo map changes.
9. Frequency by frequency
A great spatial mix is not only about where sounds go. It is about which frequencies carry direction and which ones stay solid.
| Frequency | What the engine does there |
|---|---|
| 80 or 120 Hz | Low-frequency crossover settings, so the subwoofer content stays clean and powerful |
| 200 to 300 Hz | Bass anchor filters in the binaural surround renderer, keeping weight stable around the head |
| 200 Hz, 1 kHz, 4 kHz | A 3-band EQ on every stem |
| 60 Hz to 4 kHz | The harmonic listening window used to detect key and chords |
| 4 kHz | Presence, and the high-pass that feeds binaural direction cues |
| 5 kHz | Side high-shelf for width and air on the sides |
| 8 kHz | A per-stem air filter |
| 10 kHz | High shelf for sparkle |
The philosophy is a practical one, not a law of hearing. Our ears use different cues at different frequencies: timing differences between the ears matter a lot in the lows and low mids, level differences and the shape of the outer ear matter more in the highs, and height cues live mostly up top. The engine keeps the weight of the mix stable so it never loses its punch, and lets moving sources carry their direction through the range where motion is easiest to follow. Two different things are worth separating too: the LFE channel is a dedicated effects channel in a delivery format, while bass management is how a playback system sends low frequencies to the subwoofer. The engine handles them separately, and every setting is a starting point you can change.
By hand: every one of these decisions is possible in a DAW. Keeping all of them consistent per stem while sources move through the sphere is where human time runs out.
10. A mixing partner you can talk to
Spatial9 includes a conversational mixing assistant. You talk to your mix in plain language and it acts, across 27 kinds of actions: levels, mute and solo, transport, tempo and genre, spatial mode, head tracking, stereo versus spatial comparison, ducking, masking control, harmonic movement, hearing compensation, a creative flux mode, exports and more. It can also apply lyric-driven choreography straight from the conversation.
By hand: every one of these actions exists somewhere in a traditional session, behind menus, shortcuts and plugin windows. The assistant does not replace an engineer's judgement. It makes the controls accessible in plain words, so an artist without technical training can steer their own mix.
11. One mix, up to 33 export presets, with honest rules
A spatial mix is only valuable if it reaches people. Our export system offers up to 33 presets: stereo, 5.1, 7.1, 7.1.4, 9.1.6, NHK 22.2, Auro-3D, MPEG-H 3D Audio, EBU ADM BWF, Atmos ADM BWF in bed-only, bed-plus-objects and strict modes, Ambisonics up to 9th order (100 channels), binaural, Eclipsa Audio (IAMF), game audio middleware, live-sound systems, spatial MIDI and stems with metadata.
By hand: each format usually means a separate renderer, session and round of checks. Delivering a dozen from one mix is a week of busywork that adds nothing creative.
For background on these formats, read why open source matters for immersive audio, Eclipsa Audio vs Dolby Atmos, what an ADM BWF file is and best practices for exporting stems.
The scorecard
| Feature | A. By hand | B. With today's plugins | C. Spatial9 |
|---|---|---|---|
| Movement library | Each path drawn point by point | Procedural generators and presets | 1,500+ movements on 110+ core shapes, beat-synced, phase-offset per stem |
| Section-aware staging | Planned and automated by the engineer | Partly helped by templates; coordination stays manual | Nine section profiles, from intro to outro |
| Harmony shaping the room | Chord by chord, rarely done | Still manual in our model | Breathing, key-change warps, consonance, compass |
| Lyrics choreographing space | Transcribe, time and draw every cue | Still manual in our model | 154 words and phrases, 22 actions, each one approved by you |
| Perception-inspired effects | Custom setups, one effect at a time | Some plugin equivalents, used separately | 61 effects in 9 families |
| Movement from a sentence | Not part of a classic workflow | Prompt tools, some multi-object | Across the whole mix, linked to the song, refined in conversation |
| Beat-locked transitions | Tedious and fragile | Tempo-synced panners help | Anticipation, crossover and settle on the grid |
| Frequency-aware placement | Possible with care | Possible with care | Applied per stem by default, adjustable |
| Exports | Renderer and checks per format | Fewer tools, still checks per format | Up to 33 presets from one export flow; destination rules still apply |
What this does not replace
Spatial9 does not replace the artist or the engineer. It replaces the 18,000 automation values and the hours of coordination behind them. Every move the engine makes can be heard, compared with stereo, edited or rejected. The vision, the taste and the final yes still belong to people. More effects is not the same as a better mix, and sometimes the right answer is restraint.
What changes is where human time goes. Instead of a week drawing curves, you spend an afternoon making choices. Instead of leaving the sphere empty because there is no time to fill it, you get a living space that follows the song, and you shape it.
We also hold ourselves to proving it. We are measuring total human time from source to approved release, level-matched listening comparisons, delivery acceptance and cost per approved output, and we will publish real project breakdowns as they are ready.
The game has changed
For decades immersive sound has been reserved for the biggest releases, mixed by a small number of specialists in very expensive rooms. Not because only those songs deserved it, but because the work was too slow to scale.
That era is ending. When the sections, the harmony, the words and the psychology of hearing can all shape a song's space at once, every artist can release music that surrounds the listener with intention. That is the future we are building at Spatial9, and it is already playing.
Hear it for yourself in the Music Transformation demo, learn the craft in Spatial9 Academy, or start creating in Spatial9.
Frequently asked questions
Can a human engineer achieve the same result manually?
Parts of it, yes, with enough time. A skilled engineer can place and move a few objects beautifully. What is so time-consuming by hand that it rarely happens is doing everything at once: section-aware staging, chord-by-chord harmonic motion, lyric-timed moves and psychoacoustic effects across every stem, all locked to the beat.
How many movements does Spatial9 have?
The library holds more than 1,500 movements: more than 110 core shapes in 20 families, plus roughly 1,350 movements designed for specific instruments and musical traditions. You can also create new ones from a sentence or a gesture.
Does the space change between intro, verse, chorus and outro?
Yes. Each section has its own profile. The intro is compact and low, the chorus opens wide and high, the breakdown pulls in and sinks, the climax is the widest and highest point, and the outro gently closes. Changes ease in over one to four bars.
What are harmonic movements?
Subtle motions driven by the music's key and chords: a slow two-bar breathing shaped by each chord's function, brief warps when the song changes key, a room that brightens with brighter modes, and sounds that gather on consonant chords and spread on tense ones.
Are the effects just regular plugins?
Not in the usual sense. Six are traditional tools rebuilt for space. The other 55 are perception-based effects, such as phantom sources, precedence tricks, motion afterimages, metamaterial-style rooms and time-reversal reverb. Many are marked experimental in the app.
Does the lead vocal fly around the listener?
No. The lead vocal, kick and bass keep their main positions for focus and punch. Lyric meaning drives spatial actions on the vocal layers around the lead, or on all stems if you choose.
Can I release a spatial version of a stereo song on Apple Music?
Not as Dolby Atmos music. Apple requires original multitracks or stems made from them, and does not accept upmixes of a stereo release or stems extracted from one. Use original stems for those destinations, and check each platform's rules for other workflows.
Do I still have control over the mix?
Always. Every move can be compared with stereo, adjusted or rejected, and lyric-driven moves are approved one by one. The engine does the drawing; you make the decisions.
How did you calculate 70 hours for a 5-minute pop song?
We listed every task Spatial9 automates for a 16-stem, 120 BPM, five-minute pop song in 4/4 and estimated how long a skilled immersive engineer needs to match each one, first fully by hand (about 70 hours) and then with today's spatial plugins (about 46 hours). These are illustrative estimates, not lab measurements. The Spatial9 figure of about 1 hour 50 minutes is active human time; machine processing runs on top of it.
What ROI can I expect from Spatial9?
In our illustrative model at $135 per hour, one song costs about $9,450 by hand, about $6,260 with plugins and about $257 with Spatial9, of which $9.90 is the Spatial9 price for the entire track and the rest is your own review time. That is about 97% and 96% lower modelled cost under these assumptions. It becomes a financial return when the 44 to 68 hours released per song are billed, used for more releases or replace work you would otherwise buy in. Your rates, volume and delivery list will change the numbers.
How does Spatial9 compare with today's spatial plugins?
Plugins such as motion generators, tempo-synced and audio-reactive panners, prompt-based motion tools and upmixers each solve part of the job, and some coordinate many objects at once. Spatial9's focus is connecting those decisions to the song's sections, chords, lyrics and the role of each stem in one workflow. In our illustrative model that takes a plugin-assisted mix from about 46 hours of human time to under 2.
What is included in the total cost of ownership?
Engineer time, the room, monitoring, computer and software, revisions, exports and quality checks. On the Spatial9 side we included the one-time price of $9.90 per track and charged human review time at the same full studio rate.
Do I need a DAW, plugins or a renderer to use Spatial9?
No. Spatial9 runs in the browser or through our API, with movements, effects, rendering and up to 33 export presets built in. You only need good headphones to create and review, though a final check on speakers is still wise for a professional release. If you love your DAW, you can still export ADM BWF or stems with metadata and keep working there.
Why does AI-generated music need a tool like Spatial9?
AI music platforms create songs in minutes, without studio sessions. At about 70 hours per song by hand, or about 46 with plugins, immersive mixing could never keep up with that volume. Spatial9 makes immersive output a standard part of the AI music workflow, as it already did for more than 1,000 Pozalabs tracks.