Opus 5 Guide Track Construction Analysis Review

29 Jul 2026

Opus 5 Guide Track Construction Analysis Review

Date: 2026-07-29 Source: `why-suno-hears-your-guide-as-humming-guide-track-construction.html` Analysis type: Acoustic diagnosis + Suno workflow audit

Headline

Suno is not misreading the guide. It's reproducing it faithfully. The guide has the acoustic signature of humming, not a vocal. On top of that, the workflow mode we're using is not the one Suno documents for melody transfer. Three independent failures, all fixable.

The Three Failures

1. Acoustic: The guide IS humming, not a vocal

Measured against a real lead vocal stem (from Suno's own output of this song), the guide is short on every single acoustic dimension that separates a voice from a pad - by factors of 1.9× to 206×.

| Metric | Guide v2 | Real vocal | Deficit | |--------|----------|------------|---------| | >5 kHz energy (breath, sibilance) | 0.009% | 1.852% | 206× | | Spectral flatness (noise content) | 0.00030 | 0.01616 | 54× | | Noise fraction (consonants, breath) | 0.16% | 1.90% | 12× | | 2-5 kHz presence (intelligibility) | 0.126% | 1.329% | 10.5× | | Vibrato depth | 0.023 st | 0.171 st | 7.4× | | Formant movement (vowel articulation) | 196 Hz sd | 1467 Hz sd | 7.5× | | Spectral centroid (brightness) | 917 Hz | 2650 Hz | 2.9× | | Onsets/second (syllabic rate) | 0.92 | 1.91 | 2.1× | | Voiced fraction | 91.6% | 81.0% | inverted | | Longest pitched run | 7.15 s | 6.41 s | inverted |

The guide is more continuously pitched than a human being and holds unbroken pitched tones longer than the real singer ever does. Nothing in it ever stops for a consonant or a breath. That is not a lead vocal - it is definitionally a drone.

`voice_oohs` is one vowel, forever. A real singer's formants sweep constantly because they produce different vowels and consonants. The guide has 13% of a real vocal's formant movement.

V2 went backwards on timbre: MFCC std dropped from 12.76 to 9.73, noise from 2.59% to 0.16%, presence from 0.340% to 0.126%. We fixed the composition and regressed the timbre.

2. Workflow: Wrong mode for melody transfer

Audio Influence in Create mode is NOT a melody-transfer path. Suno documents melody preservation only in:

Audio Influence in Create is a soft reference signal with no melody guarantee.

3. Settings: Audio Influence too low

Audio Influence 32 is in the "loose inspiration" regime. Community consensus puts ~50 as already loose, with 65-85 recommended when the uploaded melody/phrasing actually matters. At 32 we're explicitly telling Suno to reinterpret rather than follow.

Opus 5's Ranked Recommendations

Option A - Sing it (HIGHEST PROBABILITY)

Phone voice memo, solo, dry, on "la" or "da" syllables or the actual lyric. 20-40s covering one verse plus the hook. Then Studio Stem Cover, or Create with Audio Influence 80. This fixes all 16 metrics at once for zero engineering effort.

Option B - Solo oboe MIDI (FALLBACK)

Melody only, oboe (GM 69) or tenor sax (GM 67) or solo violin (GM 41). No pad, no guitar. Add vibrato and articulation as specified. Second best.

Option C - Current approach (STOP)

Choir-ooh melody plus pad plus guitar at Audio Influence 32. Measured to be acoustically indistinguishable from background humming.

Ten-Step Build Guide (If Staying MIDI)

1. Source: Record yourself humming (best option) 2. Strip it: Melody only. Delete halo_pad and acoustic_guitar_steel 3. Trim: 20-40s. One verse + one hook. Delete the 12.45s silence opening 4. Patch: Oboe (GM 69) first, tenor sax (GM 67), solo violin (GM 41). Never choir/voice/pad/ensemble 5. Register: Keep melody fundamentals in 220-520 Hz (A3-C5) 6. Vibrato: 5-6 Hz at 0.15-0.25 st depth on notes >0.4s. Delay onset ~200ms 7. Articulation: Raise onsets to ~1.9/s. 60-120ms gaps between notes. Cap longest run at 3s. Voiced fraction ~80% 8. Brightness: Get real energy above 2 kHz. Target ≥1% in 2-5 kHz, non-zero above 5 kHz 9. Export: WAV 44.1kHz/16-bit, normalized to -1 dBFS (not MP3) 10. Settings: Audio Influence 75-85, Weirdness <25. Prefer Studio Stem Cover or Cover over Create

Verification Thresholds

| Check | Guide v2 now | Target | |-------|-------------|--------| | 2-5 kHz energy | 0.126% | ≥1.0% | | Energy >5 kHz | 0.009% | ≥0.5% | | Spectral centroid | 917 Hz | ≥1800 Hz | | Noise fraction | 0.16% | ≥1.0% | | Vibrato depth | 0.023 st | 0.12-0.25 st | | Onsets/second | 0.92 | 1.5-2.5 | | Voiced fraction | 91.6% | 75-85% | | Longest pitched run | 7.15 s | <3 s | | Formant movement | 196 Hz | ≥800 Hz | | Duration | 138 s | 20-40 s | | Leading silence | 12.45 s | <0.5 s |

What This Means

The guide format works - we just need to provide something that IS a vocal, not something that REPRESENTS one. A phone recording of John humming beats any SoundFont patch on all 16 metrics simultaneously.

If John won't sing: solo oboe MIDI, stripped of all backing, with vibrato and articulation added, at Audio Influence 75-85, in Cover or Studio Stem Cover mode if possible.

The `voice_oohs` patch is the worst possible choice despite being named after a voice - it's an ensemble pad with one frozen vowel.

Back to studies