Why does chord audio-to-MIDI drop inner notes

Polyphonic transcription gets harder when several notes sound together because their harmonic energy overlaps in the same recording. A converter can hear the chord shape broadly correctly and still miss one middle note or invent another pitch that nobody actually played.

The problem shows up fast on piano, guitar, pads, and sampled chords. Outer notes often survive while the quiet third, seventh, or another inner voice disappears, leaving MIDI that sounds thinner than the source even though the chord still seems recognizable.

This is where MG-Mini's Full Stack chord capture mode has a much harder job than a single-note bass or vocal pass. Several fundamentals, their overtones, their attacks, and their releases all occupy the same piece of audio at once.

Polyphonic transcription has to untangle shared harmonics​

A played note produces more than its fundamental frequency. Harmonics appear above it at related frequencies, so once several notes sound together, those harmonic ladders start landing near or on top of one another.

An A minor chord gives the detector three intended notes, but the spectrum contains far more than three peaks. Some belong to the fundamentals, some belong to upper harmonics, and some frequencies receive energy from more than one note at the same time.

This overlap is a known problem in harmonic-prior models for polyphonic pitch detection. The useful bit for a producer is simple. A converter is not reading three clean labels out of the waveform. It is deciding which combinations of overlapping spectral patterns most likely came from the notes you played.

False notes often come from this mess. A strong harmonic can resemble evidence for another fundamental, especially once distortion, amp coloration, room reflections, or effects reshape the balance between partials.

Dense voicings make the puzzle worse because notes packed into the same register leave less spectral breathing room. Moving one chord tone up an octave can sometimes produce a cleaner transcription without changing the underlying harmony, simply because the notes stop competing in quite the same frequency range.

Quiet inner voices are easy to lose​

The missing note is often not random. Inner chord tones may be played more softly than the bass and melody, especially on piano or fingered guitar where the outer voices carry the musical shape.

A transcription system working from the combined signal has less evidence for a quiet middle note when louder notes are producing overlapping harmonics around it. Lowering a confidence threshold may recover the missing tone, but it can also admit more false positives from overtones and noise.

Sustain adds another layer. Hold the pedal on a piano and old notes continue ringing while new notes arrive, so the converter has to decide which frequencies belong to fresh attacks and which belong to decaying notes from the previous harmony.

Guitar creates a similar problem differently. Strings ring across chord changes, open strings can sustain much longer than fretted notes, and distortion makes the harmonic structure denser. A clean DI or dry amp signal will usually give the transcription less clutter than a wide, reverberant guitar bus.

Do not assume the loudest-looking MIDI result is the most accurate one. A conversion packed with extra notes can feel impressively detailed while quietly getting the voicing wrong. Sparse output can be wrong too, but note count by itself tells you almost nothing.

A correct chord name can hide bad MIDI​

Chord recognition and note transcription are different tests. A system can identify the harmony as C major while producing C, G, and another C in MIDI, completely missing the E that makes the chord major.

Listen to the actual voicing instead of checking only whether the chord name still works. An inversion matters. A doubled third matters. A close cluster in the middle can matter even when deleting one note leaves enough information for your ear to guess the same chord symbol.

Cleanup goes faster if you compare one chord at a time. Start with the lowest note and highest note, then listen for the tones between them. Add genuinely missing voices before deleting every suspicious upper note, because one of those upper events may be a real doubled pitch rather than a harmonic mistake.

Pedaled piano deserves another check after the notes look right. MIDI transcription generally gives you note events, not a faithful reconstruction of the original sustain-pedal performance, so long note lengths may be doing work that the source handled with pedal resonance and sympathetic ringing.

Use the instrument you plan to keep when you make the final call. A missing inner note can disappear inside a soft pad yet become painfully obvious on a dry piano, while one false upper note can turn a clean synth chord into a brittle little cluster.

If Full Stack gives you the outer shape but drops a middle voice, manual entry is usually the sane fix. Re-running the same dense chord again and again can trade one error for another, while adding a single missing note takes seconds and preserves the performance timing you already captured.
 

Attachments

  • Why does chord audio-to-MIDI drop inner notes.webp
    Why does chord audio-to-MIDI drop inner notes.webp
    24.4 KB · Views: 1

Similar threads

Sponsored

Top