AI Spatial Audio Mixing vs Manual and Plugins: What Spatial9 Automates

AI Spatial Audio Mixing vs Manual and Plugins: What Spatial9 Automates

Luiz Zanardo (CEO & Founder) ·

It is 9:00 on a Monday. A five-minute pop single lands in your inbox: 16 stems, 120 BPM, two verses, three choruses, a bridge, a breakdown and a big final lift. The label wants an immersive version by Friday.

You open the session. The sphere around the listener is empty. Every voice, every synth and every drum is waiting for you to tell it where to live, how to move, when to rise and when to fall back.

Three engineers get the same song that morning. The first does everything by hand. The second has the best spatial plugins on the market. The third uses Spatial9.

By Friday, only one of them has spent their week listening instead of drawing. Let me show you why, one task at a time.

One 5-minute pop song, three ways to mix it in 3D

Here is the song behind every number in this article: 5 minutes, 120 BPM in 4/4, 150 bars, 600 beats, 16 stems, about 400 sung words and roughly 300 chord events. Ten stems move. The kick, the bass and the lead vocal stay anchored in their main position.

The three scenarios:

Bar chart of 13 tasks for a 5-minute pop song: about 70 hours fully by hand, about 46 hours with today's plugins, about 1 hour 50 minutes of human time with Spatial9
Illustrative model for a skilled immersive engineer. Spatial9 time is active human time; machine processing runs on top of it.
TaskA. By handB. With pluginsC. Spatial9 (human time)
Session prep, stem import and gain staging1.5 h1.5 h10 min (upload and settings)
Base placement of 16 stems in 3D1.5 h1 hAutomatic
Movement design for 10 stems over 600 beats10 h3 hAutomatic, from 1,500+ movements
Section staging across about 10 section changes6 h4 hAutomatic, nine section profiles
Beat-locking and per-stem phase offsets3 h1 hAutomatic
Harmonic motion across about 300 chord events12 h10 hAutomatic
Lyric choreography: about 400 words, about 40 cues8 h7 h15 min to approve or reject
Perception-driven effects design8 h5 hPick from 61, heard in review
Frequency-aware placement and tuning4 h3 hAutomatic
Vocal clarity and ducking2 h1.5 hAutomatic
Headphone, head-tracking and stereo A/B checks2 h2 h45 min of listening
Two rounds of revisions6 h4 h30 min of notes and tweaks
Exports for six formats6 h3 h10 min to choose formats and spot-check
Total active human timeAbout 70 hAbout 46 hAbout 1 h 50 min

Two notes on fairness. Listening checks of each exported format sit inside the 45-minute review block, and destination-specific delivery rules remain your responsibility in all three scenarios. The plugin column assumes a skilled engineer who already knows the tools well.

What today's plugins genuinely solve

Credit where it is due. Today's spatial tools are good, and scenario B is a real improvement:

What they still leave on your desk

Existing tools can generate complex motion, coordinate multiple objects, react to audio and automate parts of authoring and delivery. What they do not do in our model is connect those decisions to the song's structure, harmony, lyrics and the role of each stem in one workflow. So in scenario B, the engineer still has to:

Plugins speed up the drawing. They do not take over the decisions. Spatial9 does both, and leaves you the final word.

In real life, plenty of mixes skip the harmonic motion, the lyric choreography and the perception design, sometimes by choice and often for lack of time. Spatial9 makes that a creative decision instead of a budget decision.

The cost model: three scenarios side by side

To keep the comparison fair, every scenario uses the same rate: $100 per hour for a skilled engineer, plus $35 per hour for the room, monitoring, computer and software, so $135 per working hour. Spatial9 human time is charged at that same full rate. Scenario B also includes about $50 per track for specialised spatial plugins, assuming roughly $1,500 a year of plugin licences spread over 30 immersive songs. General music software is already inside the $35 hourly overhead, so it is not counted twice. Scenario C includes the Spatial9 price: $9.90 for the entire track, one time. Everything else in that column is the engineer's own listening and review time.

Modeled cost: one song 9,450 dollars by hand, about 6,260 with plugins, about 257 with Spatial9 at 9.90 dollars per track; a 10-song album 94,500, about 62,600 and about 2,570; 100 songs 945,000, about 626,000 and about 25,700
Illustrative model at $135 per hour. Your rates, catalog and delivery list will change the numbers.
ScaleA. By handB. With pluginsC. Spatial9
One song70 h · $9,45046 h · about $6,260About 1.8 h · about $257 ($9.90 Spatial9 + review time)
10-song album700 h · $94,500460 h · about $62,600About 18 h · about $2,570 ($99 Spatial9 + review time)
100 songs a year7,000 h · $945,0004,600 h · about $626,000About 183 h · about $25,700 ($990 Spatial9 + review time)

This is a cost comparison, not a promise of profit. So here is what the numbers actually mean, in three separate ways:

The total cost of ownership also depends on what you need to deliver. A headphone-first creator can do everything on a laptop. A professional delivery still deserves a final check in a proper speaker room. Spatial9 changes how many hours you need that room for, not whether good monitoring matters.

Do you still need a DAW, plugins and a renderer?

Here is the part of the bill most people never add up. Before an engineer can spend the first of those 70 or 46 hours, someone has to buy the tools.

A traditional immersive workflow is a stack: music software (a DAW) with immersive support, a shelf of spatial plugins for panning, upmixing and effects, a renderer and encoder for each delivery format, a room with a dozen or more speakers, acoustic treatment, a powerful computer, and the training to make all of it work together.

To be fair, not everyone needs the whole stack. Some music software already renders immersive mixes on headphones, so a creator can start on a laptop. The honest comparison depends on who you are:

Who you areTraditional stackWith Spatial9
Headphone-first creatorMusic software with immersive support plus a few spatial plugins, about $500 to $3,000Good headphones, about $200 to $500, then $9.90 per track
Professional deliveryA sample full immersive room, about $18,000 to $86,000 in year one, or rented room timeCreate and review on headphones, then book room time for the final check
Catalog or AI music platformEngineers and rooms that scale with every trackThe engine and API scale with the catalog; people review, not draw
A full professional immersive room costs about 18,000 to 86,000 dollars in year one before any engineer hour, while Spatial9 runs in the browser with movements, effects, rendering and 33 export targets built in for the cost of headphones plus 9.90 dollars per track
Typical market ranges for a small professional immersive room, rounded. A headphone-first setup costs far less on both sides, and speaker-room checks can be rented when a delivery needs them.
Sample professional room budgetTypical year-one rangeWith Spatial9
Music software with immersive support$200 to $600Optional, for engineers who want to refine further
Spatial panners, upmixers and immersive plugins$500 to $3,000Built in: 1,500+ movements and 61 effects
Renderer and encoder licences$300 to $1,500Built in: rendering and up to 33 export targets
7.1.4 speakers, subwoofers and multichannel interface$8,000 to $40,000Not needed to create; good headphones, $200 to $500
Acoustic treatment for an immersive room$5,000 to $30,000Not needed to create
Local processing power for large object sessions$3,000 to $6,000A device you already own and an internet connection; processing runs on our servers
Training and calibration$1,000 to $5,000Plain language and a guided assistant
Total before any engineer hourAbout $18,000 to $86,000About $200 to $500 for headphones, then $9.90 per track

Then the stack keeps charging you. Subscriptions renew, plugins need paid upgrades, renderers change versions, computers age out, and every new delivery format means another tool to learn. That is the $35 per hour we built into the cost model for room, monitoring and software. Renting a room is an option on both sides; the real difference is how many hours of it you need.

With Spatial9, the creative work lives in one place. Movement, effects, harmony, lyrics, rendering, checks and export all happen in the same engine. You need ears, good headphones and taste. A final listen on speakers is still a good idea for a professional release.

The new age of AI music does not have a mixing desk

This matters even more for where music is going.

A growing share of new music is now created with AI music platforms. A song can be generated, rewritten and finished in minutes, often without anyone ever opening a session. These platforms produce enormous volumes of music, and their users expect results at the same speed.

There is no DAW in that workflow. No plugin chain, no renderer, no engineer booked three weeks out. The economics of the old way simply do not work: at 70 hours per song, a catalog of 10,000 AI-generated tracks would need 700,000 engineer hours, about 350 engineers working for a full year. Even with today's best plugins, at 46 hours per song, it is still 460,000 hours.

This is the role Spatial9 was built to play. We are the spatial layer of the AI music stack: a song or its stems goes in, through the app or our API, and a living, intentional immersive version comes out, ready to export. We have already done it at scale with Pozalabs, transforming more than 1,000 AI-generated tracks into immersive audio.

For AI music, immersive sound stops being a premium add-on reserved for a few releases. It can become a feature of every track.

What about the engineers who love their tools?

They are welcome, and they stay in control. Spatial9 exports ADM BWF, stems with metadata and the major immersive formats, so a professional can take the engine's work into their own session and keep refining it. The difference is that the tools become a choice, not an entry ticket.

And one rule does not change with any stack: every destination has its own delivery requirements. Some stores, Apple Music among them, require Dolby Atmos music to come from original multitracks or stems made from them. Check the rules for where your music is going, whichever way you make it.

Who is speaking, and exactly what I am claiming

I have spent my career around people who mix for a living, and I have deep respect for what they do. So let me say this carefully, and then let me say it boldly: the immersive mix that music deserves is enormously hard to draw by hand. Not because engineers lack talent, but because a 3D mix is not one decision. It is thousands of decisions, made in sync with the beat, the sections, the harmony and the words of a song, all at once.

Let me also be precise about what I am claiming. Spatial9 did not invent spatial automation. Upmixers, audio-reactive panners, procedural motion generators, prompt-driven motion tools and full immersive authoring suites already exist, and many of them are excellent. What we built is a different approach: one connected engine where the structure, the harmony, the lyrics and the way we perceive sound all shape a song's space together, and where every decision stays in the hands of the artist.

In this article I will walk through every major feature, with the real numbers inside it, and compare each one honestly with what it takes to do the same by hand or with plugins. Some of it is merely slow by hand. Some of it is so time-consuming that, in practice, almost nobody does it.

The honest truth about mixing in 3D by hand

In stereo, a mix lives on a line between two speakers. In immersive audio, every sound lives somewhere in a sphere around the listener, and it can move. That freedom is the whole magic, and it is also the trap.

Go back to our 5-minute song. Ten moving stems, each with a position in three dimensions, is 30 lanes of automation before you touch reverb, level or EQ. One point per beat across 600 beats is 18,000 automation values. Motion generators can draw many of those points for you, and that is real help. But the hard part was never the drawing. It is deciding how each move should agree with the beat, the section, the chord and the lyric being sung at that moment.

One 5-minute song has about 18,000 automation values if drawn point by point, while Spatial9 calculates motion continuously from the music
Illustrative example. Generators reduce the drawing, not the decisions.

Some projects benefit from rich spatial choreography; others benefit from restraint, and a still arrangement can be a deliberate, beautiful choice. The problem is when stillness is forced by the clock rather than chosen. Spatial9 is designed to make either choice free of repetitive setup.

Here is how, feature by feature.

1. More than 1,500 movements, built on 110+ core shapes

Our movement library holds more than 1,500 movements. At its heart are more than 110 core movement shapes in 20 families, each with a described purpose and psychoacoustic effect. Around that core sit roughly 1,350 movements designed for specific instruments and traditions: plucked strings, winds, idiophones, drums, bowed strings, mallets, brass, free reeds and keyboards, from the Afghan rubab and Persian santur to the oud, the ocarina and the West African talking drum.

The Spatial9 movement library with more than 1,500 movements built on more than 110 core shapes in 20 families, from circular orbits and spirals to cinematic, nature and psychoacoustic moves
Every movement can lock to bars and beats, with its own phase per stem.

A taste of the core families:

Under the hood, each movement is a real mathematical path in three dimensions, not a canned pan. Paths respect a minimum distance of 0.8 metres from the listener, a soft boundary that keeps sound inside the room, and a vertical range scaled to the radius. Every stem gets its own phase, so two instruments sharing an orbit stay out of step instead of overlapping.

By hand: drawing a convincing figure-8, a Mobius strip or a falling leaf means plotting dozens of points per cycle on three lanes and keeping them in time. One stem is an afternoon. A library of more than 1,500 gestures, each tuned to the character of an instrument, is something no studio could ever build session by session.

2. Every section of the song gets its own space

A song is a story with chapters. Our engine reads the structure and gives each chapter its own size, height and kind of motion.

Field expansion and height for each song section, from a compact intro and low breakdown to a wide, high chorus, drop and climax
Lead vocal, kick and bass stay anchored. Moves land on the beat.

Inside each section, moving stems choose patterns that match the energy: spirals, figure-8s, zigzags and bounces when it is high, and breathing, floating and swaying when it is low. Atmospheric stems stay soft and wave-like. Changes between sections ease in over one to four bars, depending on the profile you choose.

Equally important is what does not move. The kick, the bass and the lead vocal stay anchored, so the song keeps its spine while the space around it breathes.

By hand: building nine section-aware states and moving every stem between them musically is a full mixing project in itself. Redoing it every time the arrangement changes is where manual immersive mixing usually gives up.

3. Harmony moves the room, subtly and constantly

This is the feature people feel before they can name it. The engine estimates harmonic features, such as key, mode and chords, and maps them into adjustable spatial behaviours. These mappings are creative interpretations, not measurements of a listener's emotional response.

The engine folds 60 Hz to 4,000 Hz into 12 pitch classes, detects key, mode and chords, and uses mode brightness from Locrian to Lydian to move the field
Section energy scales the whole field: 0.85 + 0.3 x energy.

It analyses the spectrum between 60 Hz and 4,000 Hz, folds it into the 12 pitch classes and matches the result against major and minor key profiles to find the key. It then estimates each chord, with its quality, scale degree, tension and consonance. Each mode carries a brightness, from Locrian at 0.00, through Phrygian 0.17, Minor 0.33, Dorian 0.50, Mixolydian 0.67 and Major 0.83, to Lydian at 1.00.

Then come the subtle movements, the ones that make a mix feel alive rather than animated:

All of this is deliberately small, smoothed continuously and scaled by section energy. You do not hear things jumping around. You feel the harmony in the space.

By hand: a gifted engineer feels this intuitively, but tracking every chord, cadence and modulation and translating each one into gentle three-dimensional motion across a whole session would take days. Even with the best plugins this stays manual. It is the clearest case where the time cost decides the outcome, so in practice it rarely gets done. The values the engine uses are creative choices, not laws of music: they are tuned to sound natural, and you can change or switch them off.

4. The voice leads: lyrics that choreograph the space

This is one of my favourite parts of the engine, because it connects the space to the meaning of the song.

The vocal stem is transcribed word by word, its meaning is analysed with 154 built-in words and phrases, and 22 spatial actions are placed exactly on the word
Choose vocals only or all stems. Approve, edit or reject every move.

Here is the chain:

  1. The vocal stem is transcribed word by word, with a timestamp for every word.
  2. The lyrics are read for meaning, using 154 built-in words and phrases in six families (direction 39, emotion 28, compound phrases 27, nature 23, time 20 and distance 17) plus an AI reading of the context around them.
  3. Each match becomes one of 22 spatial actions, including elevation, pan, distance, scatter, collapse, orbit, freeze, isolation, spiral, wave, sweep, dome, breathe, pulse and anchor.
  4. The action lands exactly on the word.

When the singer says "falling", the sound descends. On "alone", the vocal is isolated. On "together", everything collapses inward. On "rising up", the space lifts overhead. You choose whether the cues shape the vocal layers only (doubles, harmonies, echoes and small gestures around a centred lead) or the whole mix, and you approve, edit or reject every single move.

We made one deliberate choice: the lead vocal itself stays anchored at the front for focus and intelligibility, while the singer's words lead what happens around it. The voice becomes the conductor, not a fly buzzing around your head. Vocal-aware ducking keeps room for the singer whenever they are present.

By hand: an engineer would need to transcribe the song, time every word, decide which words carry spatial meaning, and draw a move for each one on the exact syllable, across every affected stem. For one song that is days of work. Spatial9 makes lyric-timed spatial suggestions part of the workflow, with every move open to approval, editing or rejection.

5. 61 effects inspired by how we hear

Spatial9 combines conventional processing with experimental spatial effects inspired by psychoacoustics. Their names describe creative intentions; what you perceive depends on the source, the settings, the renderer, the playback system and the listener.

The Spatial9 immersive effects library with 61 effects in 9 families: psychoacoustic illusions, physics simulation, dimensional, temporal, reactive, biological, sensory, meta and narrative, and traditional spatial-aware effects
Many of these are marked experimental in the app. Use them with intention.

The library holds 61 effects in nine families. Six are the traditional tools you know, rebuilt to be spatially aware: reverb, EQ, delay, dynamics, width and distance. The other 55 are experimental effects inspired by perception research and acoustic science, including classic psychoacoustic discoveries such as the Shepard tone and the precedence and Haas effects.

By hand: some of these illusions can be built manually by an expert with a lab-grade setup, one at a time. A few have plugin cousins, many would need custom setups, and combining them with movement, harmony and lyrics in one coordinated mix is not what a chain of separate plugins was designed to do.

6. Movement from a quick idea

You type what you imagine, in plain words, and the engine gives you a movement.

From a sentence to a movement: an idea in plain words is interpreted by GravityLLM into a parametric path with a name and category, kept inside safety limits
Seconds from sentence to sound.

Write "a leaf falling slowly, then caught by a gust" and GravityLLM, our spatial intelligence, interprets the feeling and the shape you mean. It returns a real path with formulas for x, y and z over time, a name and one of six categories: circular, linear, organic, rhythmic, atmospheric or complex. Not quite right? Keep the conversation going and refine it.

Every generated movement is held within a speed of 0.5 to 2.0, a radius of 1.5 to 4 metres and a path within plus or minus 4 on every axis. Your idea gets bold, but it never gets broken.

By hand: turning a poetic image into a smooth three-dimensional curve means sketching it, translating it into keyframes and fixing the jerks. Other tools now offer prompt-driven motion too. Our difference is that the result drops straight into a session that already understands the song's sections, harmony and lyrics.

7. Perform the movement with your hands

Sometimes words are not enough and you want to show it. Spatial9 includes hand tracking and full-body pose tracking from your camera. Draw a shape in the air or move your body, and the gesture becomes a movement you can put on any stem.

By hand: ironically, this is the most human feature we have, and a mouse and automation lane cannot capture the arc of a gesture in three dimensions.

8. Beat-locked motion, with taste built in

Movement that ignores the groove sounds like a screensaver. In Spatial9, the cycle of an orbit can be set in bars, phases offset in beats, and section changes eased over one to four bars.

Transitions follow a three-phase curve: anticipation, crossover and settle, with the crossover complete at 75% of the move. Speed is capped and every move lasts at least half a second, which keeps motion smooth rather than whipping past your ears. Pads, atmospheres and effects get an extra slow drift of a few centimetres over 16 to 64 bars, the kind of life you feel more than notice.

By hand: keeping hundreds of moves locked to tempo, with consistent anticipation and settle, is the definition of tedious, and it is the first thing to fall apart when the tempo map changes.

9. Frequency by frequency

A great spatial mix is not only about where sounds go. It is about which frequencies carry direction and which ones stay solid.

Frequency map from 20 Hz to 20 kHz showing LFE crossovers at 80 or 120 Hz, bass anchor filters at 200 to 300 Hz, a 3-band EQ at 200 Hz, 1 kHz and 4 kHz, presence and binaural high-pass at 4 kHz, side shelf at 5 kHz, air at 8 kHz and high shelf at 10 kHz
Weight stays stable and focused while the upper range carries clear direction cues. Starting points, not fixed rules.
FrequencyWhat the engine does there
80 or 120 HzLow-frequency crossover settings, so the subwoofer content stays clean and powerful
200 to 300 HzBass anchor filters in the binaural surround renderer, keeping weight stable around the head
200 Hz, 1 kHz, 4 kHzA 3-band EQ on every stem
60 Hz to 4 kHzThe harmonic listening window used to detect key and chords
4 kHzPresence, and the high-pass that feeds binaural direction cues
5 kHzSide high-shelf for width and air on the sides
8 kHzA per-stem air filter
10 kHzHigh shelf for sparkle

The philosophy is a practical one, not a law of hearing. Our ears use different cues at different frequencies: timing differences between the ears matter a lot in the lows and low mids, level differences and the shape of the outer ear matter more in the highs, and height cues live mostly up top. The engine keeps the weight of the mix stable so it never loses its punch, and lets moving sources carry their direction through the range where motion is easiest to follow. Two different things are worth separating too: the LFE channel is a dedicated effects channel in a delivery format, while bass management is how a playback system sends low frequencies to the subwoofer. The engine handles them separately, and every setting is a starting point you can change.

By hand: every one of these decisions is possible in a DAW. Keeping all of them consistent per stem while sources move through the sphere is where human time runs out.

10. A mixing partner you can talk to

Spatial9 includes a conversational mixing assistant. You talk to your mix in plain language and it acts, across 27 kinds of actions: levels, mute and solo, transport, tempo and genre, spatial mode, head tracking, stereo versus spatial comparison, ducking, masking control, harmonic movement, hearing compensation, a creative flux mode, exports and more. It can also apply lyric-driven choreography straight from the conversation.

By hand: every one of these actions exists somewhere in a traditional session, behind menus, shortcuts and plugin windows. The assistant does not replace an engineer's judgement. It makes the controls accessible in plain words, so an artist without technical training can steer their own mix.

11. One mix, up to 33 export presets, with honest rules

A spatial mix is only valuable if it reaches people. Our export system offers up to 33 presets: stereo, 5.1, 7.1, 7.1.4, 9.1.6, NHK 22.2, Auro-3D, MPEG-H 3D Audio, EBU ADM BWF, Atmos ADM BWF in bed-only, bed-plus-objects and strict modes, Ambisonics up to 9th order (100 channels), binaural, Eclipsa Audio (IAMF), game audio middleware, live-sound systems, spatial MIDI and stems with metadata.

By hand: each format usually means a separate renderer, session and round of checks. Delivering a dozen from one mix is a week of busywork that adds nothing creative.

For background on these formats, read why open source matters for immersive audio, Eclipsa Audio vs Dolby Atmos, what an ADM BWF file is and best practices for exporting stems.

The scorecard

FeatureA. By handB. With today's pluginsC. Spatial9
Movement libraryEach path drawn point by pointProcedural generators and presets1,500+ movements on 110+ core shapes, beat-synced, phase-offset per stem
Section-aware stagingPlanned and automated by the engineerPartly helped by templates; coordination stays manualNine section profiles, from intro to outro
Harmony shaping the roomChord by chord, rarely doneStill manual in our modelBreathing, key-change warps, consonance, compass
Lyrics choreographing spaceTranscribe, time and draw every cueStill manual in our model154 words and phrases, 22 actions, each one approved by you
Perception-inspired effectsCustom setups, one effect at a timeSome plugin equivalents, used separately61 effects in 9 families
Movement from a sentenceNot part of a classic workflowPrompt tools, some multi-objectAcross the whole mix, linked to the song, refined in conversation
Beat-locked transitionsTedious and fragileTempo-synced panners helpAnticipation, crossover and settle on the grid
Frequency-aware placementPossible with carePossible with careApplied per stem by default, adjustable
ExportsRenderer and checks per formatFewer tools, still checks per formatUp to 33 presets from one export flow; destination rules still apply

What this does not replace

Spatial9 does not replace the artist or the engineer. It replaces the 18,000 automation values and the hours of coordination behind them. Every move the engine makes can be heard, compared with stereo, edited or rejected. The vision, the taste and the final yes still belong to people. More effects is not the same as a better mix, and sometimes the right answer is restraint.

What changes is where human time goes. Instead of a week drawing curves, you spend an afternoon making choices. Instead of leaving the sphere empty because there is no time to fill it, you get a living space that follows the song, and you shape it.

We also hold ourselves to proving it. We are measuring total human time from source to approved release, level-matched listening comparisons, delivery acceptance and cost per approved output, and we will publish real project breakdowns as they are ready.

The game has changed

For decades immersive sound has been reserved for the biggest releases, mixed by a small number of specialists in very expensive rooms. Not because only those songs deserved it, but because the work was too slow to scale.

That era is ending. When the sections, the harmony, the words and the psychology of hearing can all shape a song's space at once, every artist can release music that surrounds the listener with intention. That is the future we are building at Spatial9, and it is already playing.

Hear it for yourself in the Music Transformation demo, learn the craft in Spatial9 Academy, or start creating in Spatial9.

Frequently asked questions

Can a human engineer achieve the same result manually?

Parts of it, yes, with enough time. A skilled engineer can place and move a few objects beautifully. What is so time-consuming by hand that it rarely happens is doing everything at once: section-aware staging, chord-by-chord harmonic motion, lyric-timed moves and psychoacoustic effects across every stem, all locked to the beat.

How many movements does Spatial9 have?

The library holds more than 1,500 movements: more than 110 core shapes in 20 families, plus roughly 1,350 movements designed for specific instruments and musical traditions. You can also create new ones from a sentence or a gesture.

Does the space change between intro, verse, chorus and outro?

Yes. Each section has its own profile. The intro is compact and low, the chorus opens wide and high, the breakdown pulls in and sinks, the climax is the widest and highest point, and the outro gently closes. Changes ease in over one to four bars.

What are harmonic movements?

Subtle motions driven by the music's key and chords: a slow two-bar breathing shaped by each chord's function, brief warps when the song changes key, a room that brightens with brighter modes, and sounds that gather on consonant chords and spread on tense ones.

Are the effects just regular plugins?

Not in the usual sense. Six are traditional tools rebuilt for space. The other 55 are perception-based effects, such as phantom sources, precedence tricks, motion afterimages, metamaterial-style rooms and time-reversal reverb. Many are marked experimental in the app.

Does the lead vocal fly around the listener?

No. The lead vocal, kick and bass keep their main positions for focus and punch. Lyric meaning drives spatial actions on the vocal layers around the lead, or on all stems if you choose.

Can I release a spatial version of a stereo song on Apple Music?

Not as Dolby Atmos music. Apple requires original multitracks or stems made from them, and does not accept upmixes of a stereo release or stems extracted from one. Use original stems for those destinations, and check each platform's rules for other workflows.

Do I still have control over the mix?

Always. Every move can be compared with stereo, adjusted or rejected, and lyric-driven moves are approved one by one. The engine does the drawing; you make the decisions.

How did you calculate 70 hours for a 5-minute pop song?

We listed every task Spatial9 automates for a 16-stem, 120 BPM, five-minute pop song in 4/4 and estimated how long a skilled immersive engineer needs to match each one, first fully by hand (about 70 hours) and then with today's spatial plugins (about 46 hours). These are illustrative estimates, not lab measurements. The Spatial9 figure of about 1 hour 50 minutes is active human time; machine processing runs on top of it.

What ROI can I expect from Spatial9?

In our illustrative model at $135 per hour, one song costs about $9,450 by hand, about $6,260 with plugins and about $257 with Spatial9, of which $9.90 is the Spatial9 price for the entire track and the rest is your own review time. That is about 97% and 96% lower modelled cost under these assumptions. It becomes a financial return when the 44 to 68 hours released per song are billed, used for more releases or replace work you would otherwise buy in. Your rates, volume and delivery list will change the numbers.

How does Spatial9 compare with today's spatial plugins?

Plugins such as motion generators, tempo-synced and audio-reactive panners, prompt-based motion tools and upmixers each solve part of the job, and some coordinate many objects at once. Spatial9's focus is connecting those decisions to the song's sections, chords, lyrics and the role of each stem in one workflow. In our illustrative model that takes a plugin-assisted mix from about 46 hours of human time to under 2.

What is included in the total cost of ownership?

Engineer time, the room, monitoring, computer and software, revisions, exports and quality checks. On the Spatial9 side we included the one-time price of $9.90 per track and charged human review time at the same full studio rate.

Do I need a DAW, plugins or a renderer to use Spatial9?

No. Spatial9 runs in the browser or through our API, with movements, effects, rendering and up to 33 export presets built in. You only need good headphones to create and review, though a final check on speakers is still wise for a professional release. If you love your DAW, you can still export ADM BWF or stems with metadata and keep working there.

Why does AI-generated music need a tool like Spatial9?

AI music platforms create songs in minutes, without studio sessions. At about 70 hours per song by hand, or about 46 with plugins, immersive mixing could never keep up with that volume. Spatial9 makes immersive output a standard part of the AI music workflow, as it already did for more than 1,000 Pozalabs tracks.

Try Spatial9 free