Beds, objects, and ambisonics do different jobs

The Dolby Atmos Renderer takes 128 inputs in total, and every bed channel eats one, so a 7.1.2 bed costs ten objects before you place a sound. That trade is the whole argument. These three approaches are not quality tiers stacked on top of each other; they are different ways of storing where a sound lives.

A channel is the oldest and simplest of them. It is one stream of audio meant to play out of one speaker in a known position, carrying no information about where that position is, which is why 7.1.4 is just twelve channels feeding twelve specific boxes. Nothing thinks at playback. Nothing adapts.

That simplicity is also the weakness, because the mix only lands correctly in a room whose speaker map matches the studio's. Beds are where host limits bite hardest, since a bed is nothing more than a wide bus, and the number of channels a host will pass decides whether you get a real 7.1.2 or a stack of stereo pairs doing an impression of one.

Objects buy precision at a rendering cost​

An object is a single sound plus positional metadata that changes over time. No speaker assignment travels with it. The renderer at the other end reads the coordinates and works out which of the available speakers should play how much of that sound, which is why one master serves headphones, a soundbar, a 5.1 system, and a cinema without three separate mixes being made.

Dolby Atmos runs a hybrid of both, pairing a bed of up to 7.1.2 with as many as 118 objects. The bed handles anything diffuse and static. Objects handle anything that needs to be in a specific place or needs to move, which in music means most of the interesting material. Dolby is not alone in this. DTS and MPEG-H both carry object metadata as well, and MPEG-H can hold channel, object, and scene-based elements in a single stream.

Music mixes can skip the bed entirely and run everything as objects, and plenty do. There is a catch worth knowing before you spend a week placing 90 elements. Spatial coding at delivery folds roughly 128 elements down to somewhere near 12 to 16 perceptually distinct clusters, so the count you author is not the count that reaches a listener. Objects sitting in near-identical positions get merged whether you like it or not.

Ambisonics describes the field, not the speakers​

Scene-based audio takes a third route. Rather than storing speakers or sound sources, it stores the entire sound field at one point in space as a set of mathematical layers, with directional information encoded independently of any speaker layout.

The channel count follows a formula, which is order plus one, squared. First order gives four channels, labeled W, X, Y, and Z. Second order gives nine. Third order gives sixteen, and seventh order gives sixty-four. Precision climbs with the order, and first order alone is fairly crude about direction, which is why serious work reaches for higher orders and accepts the channel count that comes with them.

A sound field stored as mathematical layers has one property the other two cannot match. Rotating it is a single cheap operation applied to the whole field at once. Turn your head in a VR headset and an object-based mix has to re-render every object individually, while an ambisonic scene just rotates. That is the entire reason the format owns 360 video.

Your delivery target picks the format for you​

YouTube accepts first-order ambisonics for 360 and VR video, at 48 kHz, in AmbiX format with ACN channel ordering and SN3D normalization. Four channels for the plain version. Six if you want head-locked stereo riding alongside, meaning music or narration that stays put while the viewer looks around.

The container rules are specific. PCM in a MOV file works for both. AAC in MP4 or MOV needs at least 256 kbps for the four-channel version, and OPUS in MP4 requires channel mapping family 2 at 512 kbps, rising to 768 kbps once head-locked stereo is added.

Music streaming wants none of that. It wants Atmos, which means beds and objects, and consumer spatial audio systems do not decode ambisonics directly at all. An ambisonic master gets converted to a channel bed before anyone hears it, so treating ambisonics as a capture and production format on its way to becoming something else is closer to how it actually gets used.
 

Attachments

  • Beds, objects, and ambisonics do different jobs.webp
    Beds, objects, and ambisonics do different jobs.webp
    172.9 KB · Views: 1

Similar threads

Sponsored

Top