Channels, Objects, Scenes and Binaural: The Language of Immersive Formats

Channels, Objects, Scenes and Binaural: The Language of Immersive Formats

Spatial9 Team ·

Lesson 3 of 14 in Spatial9 Academy, module 1: Foundations. Practice in the demo: Music Transformation.

Every language has a grammar. Immersive audio does too. When a producer asks for "a 7.1.4 bed with objects", a VR team asks for "third-order Ambisonics", or a distributor asks for "an ADM BWF master", they are speaking that language.

Once you can speak it, a lot of doors open. You can read a delivery spec without panic. You can talk to film mixers, game audio teams, streaming platforms and developers as an equal. You can choose the right format for each project instead of guessing.

The good news is that the grammar is simpler than it looks. There are only four big ideas. Everything else is detail.

What you will learn

The four paradigms at a glance

Every immersive format answers one question: how do we write down space so a playback system can recreate it? There are four main answers.

Four cards comparing channel-based, object-based, scene-based and binaural audio, each with an icon and what it stores
Four ways to write down space: speaker feeds, sounds with positions, a full sound field, or two ear signals.

Think of them like this. Channel-based audio is a set of finished speaker feeds. Object-based audio is a set of sounds plus instructions about where they go. Scene-based audio is a recording of the whole sound field around a point. Binaural audio is what arrives at two ears, ready for headphones.

Real productions mix these ideas. A cinema mix might combine a channel bed for ambience with objects for moving sounds. A VR project might store a scene and render it to binaural on playback. Knowing all four lets you understand any of these combinations.

Channel-based audio and the 7.1.4 notation

Channel-based audio is the oldest approach. Each channel is meant for one speaker in a known position. Stereo is channel-based: one feed for the left speaker, one for the right.

Surround and immersive layouts use a simple three-part notation: main.LFE.height.

So 5.1 means five main channels and one LFE channel. 7.1.4 means seven main channels, one LFE channel and four height channels, for twelve channels in total.

A 7.1.4 speaker layout seen from above with seven ear-level speakers, a subwoofer and four height speakers
A 7.1.4 layout: seven speakers at ear level, one LFE channel and four height speakers overhead.

The strength of channel-based audio is simplicity. What you hear in the studio is what the speakers play. The weakness is rigidity. A 7.1.4 mix is designed for 7.1.4. Playing it on a different layout requires downmixing or upmixing, and the result is never quite what you intended.

Object-based audio: sounds with coordinates

Object-based audio changes the question. Instead of saying "play this in the left surround speaker", you say "this sound is here", with coordinates such as azimuth, elevation and distance, which you met in How We Hear in 3D.

Each audio object is an audio signal plus metadata: position, size, gain and how these change over time. A renderer reads that metadata on playback and works out how to feed the speakers actually present, whether that is a cinema, a 7.1.4 studio, a soundbar or headphones.

This is why objects are so powerful. One mix can adapt to many rooms. A bird that flies overhead is stored as a bird with a path, and each playback system draws that path as well as it can.

Most object-based workflows also use beds: channel-based layers for sounds that do not need to move, such as ambience or a wide pad. Beds plus objects is a very common structure in immersive music and film.

Scene-based audio: Ambisonics

Ambisonics takes a third approach. It does not store speaker feeds or individual sounds. It stores the whole sound field around a listening point, using a set of spherical harmonic components.

First-order Ambisonics uses four channels. Higher orders add more channels and more spatial precision. The rule is simple: order n uses (n+1) squared channels.

Ambisonic orderChannelsTypical use
First order4360 video, field recording with Ambisonic microphones
Second order9More precise scenes
Third order16Detailed immersive scenes and production work

The magic of Ambisonics is rotation. Because the whole sphere is stored, a player can rotate the scene to match a viewer's head or a VR camera. That is why Ambisonics became the natural choice for 360 video and virtual reality. For the full story, read Ambisonics explained.

Binaural: the headphone destination

Binaural is not a separate way of creating space so much as a way of delivering it to two ears. A binaural renderer takes channels, objects or a scene and applies head-related transfer functions (HRTFs) to produce a left and a right signal with 3D cues baked in.

The result is a normal two-channel file that plays on any headphones. That makes binaural the most universal destination for immersive audio today. Its limitation is that it is designed for headphones. Played over speakers, the 3D cues mostly collapse. For a deeper comparison, read 8D vs binaural vs spatial audio.

The deliverables you will actually export

Paradigms are ideas. Deliverables are files. Here are the ones you will meet most often.

ADM BWF. The Audio Definition Model (ADM) is an ITU standard for describing immersive audio metadata: channels, objects, scenes and how they fit together. An ADM BWF file stores that metadata inside a Broadcast Wave Format file alongside the audio. It is a widely used interchange master for object-based mixes. Learn more in What is an ADM BWF file?.

Eclipsa Audio (IAMF). IAMF is an open, royalty-free format from the Alliance for Open Media, with version 1.0 published in 2023. It can carry channel-based layers or Ambisonic scenes, and it supports the codecs Opus, AAC-LC, FLAC and LPCM. YouTube has accepted Eclipsa Audio since January 2025, and Samsung's 2025 TVs and soundbars play it.

Dolby Atmos. A licensed, object-based format that supports up to 128 channels of beds and objects. It is common in cinema and in music streaming services that offer immersive music. For a detailed comparison with Eclipsa Audio, read Eclipsa Audio vs Dolby Atmos.

Binaural and Ambisonics files. Binaural stereo for headphones and social platforms. Ambisonics for VR and 360 video.

At Spatial9, deliverables depend on the workflow and the plan, and include ADM, Eclipsa Audio (IAMF), binaural and Ambisonics.

Which format for which destination

The best way to choose a format is to start at the end: where will people hear this?

Decision map linking YouTube to Eclipsa Audio, streaming services with Atmos to an ADM BWF master, headphones and social to binaural, and VR and 360 video to Ambisonics
Start with the destination, then choose the deliverable.

Platforms update their specifications over time, so always check the current delivery requirements before export. For a destination-by-destination guide, see where to publish spatial audio.

Try it in the demo

Wear headphones. The demo shows the notice "Best Experienced with Headphones", and the difference between formats is subtle on speakers. Music Transformation carries the badge Live.

  1. Open /demo/music and pick the tile "3D Spatial Audio" under "Choose a mode".
  2. In "Master Controls", press play and set "Volume" to a comfortable level. Make sure the "Spatial Audio" toggle is on.
  3. Under "Spatial Format", select "Spatial Sphere". The demo describes it as Higher-Order Ambisonics (HOA): the scene-based paradigm.
  4. Listen for 30 seconds. Focus on how evenly the sound wraps around you and how smoothly sounds sit between positions.
  5. Switch to "Immersive Mix (ADM)", described as Channel and Object-based spatial audio. Listen for the same 30 seconds.
  6. Compare the two. Note the clarity of individual stems such as "Synth Lead" and "Pad" in the "Stems (23)" list, and how precise their positions feel in each mode.
  7. Select one stem, such as "Riser", mute the others with the speaker icons, and switch formats again. A single source makes the difference between scene and object rendering easier to hear.

Key terms

TermMeaning
Channel-basedAudio stored as feeds for specific speakers in a fixed layout.
7.1.4Seven main channels, one LFE channel and four height channels.
Audio objectAn audio signal plus position metadata, rendered to any layout.
BedA channel-based layer, often used for static or ambient sounds alongside objects.
AmbisonicsScene-based audio that stores a full sphere; order n needs (n+1) squared channels.
BinauralTwo-channel headphone audio carrying 3D cues from HRTF processing.
ADM BWFA Broadcast Wave file containing Audio Definition Model metadata for immersive audio.
IAMFThe open, royalty-free immersive format from the Alliance for Open Media, known as Eclipsa Audio.

Check your understanding

  1. How many channels are in a 7.1.4 layout, and what does each number mean?
  2. What makes an object-based mix more flexible than a channel-based mix?
  3. How many channels does third-order Ambisonics use, and why?
  4. Why is binaural the most universal destination for immersive audio?
  5. Which deliverable would you choose for an immersive music video on YouTube?

Answers

Assignment

Write a one-page "delivery plan" for a fictional release: a four-minute single with a music video, a vertical social clip and a 360 live-session video. For each destination, name the deliverable format, the paradigm it uses (channel, object, scene or binaural), and one risk to check before delivery. Include a simple table and a short paragraph explaining which master you would create first and why. Deliver it as a PDF.

Frequently asked questions

What is the difference between channel-based and object-based audio?

Channel-based audio stores a finished signal for each speaker in a fixed layout, while object-based audio stores individual sounds plus position metadata, so a renderer can place them correctly on whatever speakers or headphones are available.

What does 7.1.4 mean in immersive audio?

The notation 7.1.4 means seven main speakers around the listener, one low-frequency effects channel and four height speakers overhead, for a total of twelve channels.

How many channels does Ambisonics use?

Ambisonics uses (n+1) squared channels for order n, so first order uses 4 channels, second order 9 and third order 16.

What is the best format for spatial audio on YouTube?

YouTube has accepted Eclipsa Audio, the consumer name of the open IAMF format, since January 2025, which makes it the natural choice for immersive audio on YouTube.

Dolby and Dolby Atmos are trademarks of Dolby Laboratories. Spatial9 is not affiliated with or endorsed by Dolby Laboratories.

Next step

You can now speak the language of immersive formats. In the next lesson, The Spatial9 Approach: AI as a Creative Partner in Immersive Production, you will see how AI and human taste work together to create these deliverables. Visit the Spatial9 Academy course home to see every lesson.

Try Spatial9 free