Spectral selection lifts one sound out of a mix

At 48 kHz, a 2048-sample analysis window resolves about 23 Hz per bin but smears anything faster than 43 milliseconds into a single blur.

That number is the whole game. A spectral edit is a rectangle drawn in time and frequency, and the grid you are drawing on was fixed the moment the window size was set.

The reason you are doing this at all is that a stereo capture of everything your desktop played arrives pre-mixed, with the browser tab, the standalone synth, and the video call already summed into one pair. Nothing downstream can unmix it. What you can do is reach into the picture and take a region out.

The window setting decides what you can separate​

Small windows see time clearly and frequency badly.

A 256-sample window at 48 kHz covers 5.3 milliseconds and gives you 187 Hz per bin. Audacity's own demonstration shows this cleanly, with two closely spaced clicks visible as two events at a window of 256 and merged into one at 2048. Push the window to 4096, and you get 11.7 Hz resolution across 85 milliseconds, which is the opposite failure.

Pitch decides which failure hurts. A semitone at 82 Hz is a gap of under 5 Hz, so 187 Hz bins cannot tell a low E from anything near it. The same semitone up at 1760 Hz is 105 Hz wide, comfortably resolved by a small window with time resolution to spare.

So the setting follows the target, not the session. Chasing a click or a mouth noise, go short. Chasing a sustained hum or a bass note sitting under everything, go long and accept that fast events in the selection will spread.

Get this wrong, and the edit fails before you draw anything. You will select a rectangle that looks right on screen and contains three times the material you meant.

Reconstruction mode decides how the edit sounds​

Once the region is gone, whatever fills it has to be rebuilt, and the filter used to rebuild it has a signature.

Zero-phase and linear-phase processing spread filter ringing evenly on both sides of an event. Minimum-phase processing puts all of it after. That is the entire practical difference, and it decides which artifacts you hear.

Your ears are asymmetric about this. Forward masking is considerably stronger than backward masking, so energy arriving just after a transient hides underneath it, while ringing that arrives before the transient has nothing covering it. Listeners tend to describe the zero-phase version of a click as having a chirp on the front.

Percussive material therefore wants minimum phase. Sustained tonal material is the other way round, because the phase relationships between partials hold the timbre together and a nonlinear group delay will shift them against each other.

Neither choice is free. Linear phase costs you latency and a pre-echo you might notice on a bare hit. Minimum phase keeps the transient honest and moves your partials around.

Some sounds share a bin and cannot be split​

Masking has a hard floor. If two sources put energy into the same bin at the same instant, no selection can tell them apart, and you attenuate both or neither.

Repair tools work around this by borrowing instead of subtracting. Attenuate pulls magnitudes in the selection down to match the surrounding area without resynthesis. Replace interpolates a new region from the audio either side of it. Pattern hunts for similar material nearby and copies it in. Partials and Noise track harmonic components across the gap, link them, then handle the leftover noise separately, which is why it survives vibrato that defeats the simpler modes.

The borrowing has documented limits. In iZotope RX, horizontal Attenuate and Replace stop at ten seconds of selection, Pattern and Partials plus Noise at four, and only vertical Attenuate is unbounded. Longer selections get quietly downgraded to a mode that can handle them.

Direction matters as much as mode. Horizontal interpolation reads from before and after your selection, which works on steady material and falls apart on anything that changes fast across the gap. Vertical reads from above and below, which is fine in a sparse spectrum and useless where partials are packed. The 2D option takes from everywhere and produces the mushiest result of the three, which is sometimes exactly what you want on a short selection nobody will inspect closely.
 

Attachments

  • Spectral selection lifts one sound out of a mix.webp
    Spectral selection lifts one sound out of a mix.webp
    172.9 KB · Views: 1

Sponsored

Top