Thread Context
Thread Context
Key facts and notes for this thread. Updated by agent, survives context compaction.
Numbers & Values
FILES: A = "The Rinse Light -A.mp3" (231.55s / 3:51). B = "The Rinse Light -B.mp3" (208.99s / 3:29). B is 22.6s shorter. Confirmed genuinely different audio (zero-lag corr -0.0007).
TEMPO: both share fundamental 69.84 BPM (autocorrelation strongest peak both files). B preserved A's tempo. Beat salience stronger in B (19171 vs 16896). Beat-interval CV: A 0.0551, B 0.0573. Tempo drift across song: A +1.05 BPM (speeds up), B -2.74 BPM (relaxes).
KEY: A reads C major (0.913) / A minor (0.825). B reads F major (0.920) / D minor (0.673). Tonal centre shifted.
LOUDNESS: both normalised to -14.0 LUFS integrated, both LRA 7.1 LU. Crest factor A 16.52 dB, B 15.22 dB. Dynamic span p5-p95: A 21.66 dB, B 21.97 dB. Macro dynamic range essentially IDENTICAL - B did not expand it.
STEREO: A side/mid 0.384, L/R corr 0.743 (wider). B side/mid 0.252, L/R corr 0.881 (narrower/more centred). A is the wider mix overall.
VERSE-TO-CHORUS LIFT: A +0.70 dB, centroid -345 Hz, sub-bass +7.9 pts, onsets +15%. B +1.42 dB, centroid -78 Hz, sub-bass +16.0 pts, onsets +1%. Both weak on level; A roughly half of B.
SECTION LEVEL ARC (RMS dB rel peak): A: intro -24.87, V1 -22.27, Ch1a -19.55, Ch1b -18.80, V2 -20.08, Pre -18.07, Ch2 -18.24, Bridge -20.81, Ch3+outro -21.17. Peak at Pre/Ch2; final chorus 2.9 dB QUIETER than Ch2. B: intro -28.91, V1 -22.06, Ch1 -17.60, V2 -17.01, Ch2 -16.55, Bridge -17.98, Ch3+outro -20.18. Monotonic build to Ch2 then release.
SUB-BASS % BY SECTION: A: 1.2, 6.2, 7.8, 10.4, 7.7, 4.6, 15.9, 18.7, 22.0 B: 2.1, 4.9, 13.1, 10.0, 31.3, 30.7, 26.1 (B commits ~2x sub in choruses)
STEREO WIDTH BY SECTION: A 0.212 to 0.521 (progressive widening, ends widest). B 0.192 to 0.320.
PERCUSSIVE RATIO: A bridge 17.1%. B bridge 27.5% (B makes bridge a rhythmic event).
ONSET MICRO-TIMING vs 16th grid: A mean abs dev 13.4 ms, std 20.7 ms. B 11.7 ms, std 18.7 ms. B tighter.
SPECTRAL: centroid A 1802 Hz, B 1896 Hz. Rolloff85 A 3954, B 4159. Flatness A 0.00972, B 0.01199.
TRANSCRIPT EMOTION TAGS: A = sad on all 8 segments (uniform). B = neutral/sad/neutral/sad/sad/neutral (differentiates verse vs chorus).
STRUCTURE: A has 9 sections incl. split chorus halves (Ch1a 57-71, Ch1b 71-97) and separate Pre/V2b (124-137). B merged chorus halves into single 33s chorus and merged V2+pre into one 35s verse. Same lyric, same 281 word count.
LYRIC VARIANTS: A "like I forgot to close" vs B "like it forgot to close". A "I lined the coins" vs B "I line the coins". A "the amber word rinse" vs B transcribed "amber wood rinse". A "worse prayers" vs B "worst prayers" (B diction less clear on those two).
=== SONG 3: THE GUIDE TRACK ITSELF (guide-track-upload-safe.mp3) === 196.03s (3:16). -18.6 LUFS, LRA 6.3 LU, mean -22.4 dB, max -1.6 dB. L/R corr 0.956, side/mid 0.180 (near-mono). 72 BPM -> bar 3.3333s, beat 0.8333s, 58.8 bars total.
THREE LAYERS IDENTIFIED (MP3 render, so levels are amplitude not MIDI velocity): 1. vocal_marker = melody. 105 notes, range C3-D4 (14 st), median G3, median dur 1.300s (1.56 beats), 78% >0.5s, 57% >1s. Each note preceded by a ~23-35ms portamento glide 1 semitone below (70% of 66 short mid events are exactly +1 st below the following melody note, median dt -23ms). Median CQT peak -5.86 dB. 2. room_hum = SUSTAINED MOVING BASS LINE, not a static hum. Roots A2 x7, G2 x7, F2 x3, D2 x1, E2 x1. Durations 2 bars (6.6s) or 4 bars (13.35s). Median CQT peak -3.75 dB = LOUDEST layer. 3. muted_pulse = quiet high clicks C#6 (43), B5 (42), G5 (10). Median peak -32.43 dB. IOI median 0.406s (~8th note). MONOPHONIC: 166 of 169 groups are single notes. NO CHORD STABS ANYWHERE.
BALANCE: melody minus bass = -2.11 dB (melody is BELOW the loudest backing element). Melody minus pulse = +26.57 dB. In its own band (126-300Hz) melody is +28.94 dB above bass harmonic bleed, so it is clearly audible - not masked, but not dominant either.
STRUCTURE (derived from melodic repeats + bass harmony, mutually confirming): bars 1-7 (0-22.7s) opening. NO INSTRUMENTAL INTRO - first melody note at t=0.000, downbeat of bar 1. bars 8-13 (24.5-42.0) VERSE 1 = ALPHA motif, 11 notes, over 4-bar A2 pedal bars 14-16 (45.4-52.0) CHORUS 1 = BETA motif "A3 G3 D3" + long G3 (2.87s), over F2(2 bars)->G2(2 bars) bars 17-20 (53.3-65.2) TAG 1 bars 21-24 (66.7-76.1) link bars 25-29 (77.6-95.7) VERSE 2 = ALPHA-2 over 4-bar A2 pedal bars 30-32 (99.0-105.7) CHORUS 2 = BETA-2 over F2->G2 bars 33-36 (107-119) TAG 2 (contains the single D4) bars 37-38 (120.4-125.1) BRIDGE, E2 bass = E major, contains the ONLY G#3 in the track bars 39-41 (127-135.6) bridge continuation bars 42-43 (137.4-143) CHORUS 3 variant bars 44-45 (144.5-149.7) link bars 46-51 (151.7-169.2) VERSE 3 = ALPHA-3 over 4-bar A2 pedal bars 52-54 (172.6-179.1) CHORUS 4 = BETA-4 bars 55-59 (180.5-189) outro The 4-bar A2 pedal aligns exactly with each ALPHA; F2->G2 2-bar changes align exactly with each BETA. Mutual confirmation of the map.
Q1 HOOK APEX: BETA apex = A3 on ALL FOUR statements. Contour self-similarity r = +1.000, spread 0.00 st (pitch sequence A3-G3-D3 literally identical each time). BUT A3 is only the 10th-highest pitch used; the track reaches C4 8x and D4 1x, all OUTSIDE the chorus. So the apex is perfectly consistent but 3 semitones below the track's own ceiling and identical to the verse's apex. ALPHA (verse) apex: A3, B3, A3 - varies by 2 st.
Q2 RANGE: VERSE (ALPHA) span 9, 11, 9 st. CHORUS (BETA) span 7, 7, 7, 7 st. Chorus is NARROWER than verse by ~2-4 st. Guide DOES encode chorus contraction. Global range C3-D4 = 14 st. D4 occurs exactly ONCE (bar 33, in a tag).
Q3 HARMONIC RHYTHM: verse = 4-bar A minor pedal = 0.25 chords/bar. Chorus = F 2 bars then G 2 bars = 0.5 chords/bar. CHORUS IS 2x FASTER THAN VERSE - inverted differentiation. Also the chorus never touches the tonic (F->G plagal, unresolved).
Q4 BALANCE: see above. Vocal 2.11 dB BELOW bass, 26.57 dB above pulse.
Q5 INTRO: ZERO. First melody note at t=0.000s on the downbeat of bar 1.
Q6 BRIDGE (bars 37-38): notes/beat 0.48 vs neighbours 0.38 and 0.42 -> density INCREASES. rest 12% = LOWEST of any section. medDur 1.68s = longest. span 4 st = narrowest. So sustained and narrow but NOT sparse; no dropout encoded. A 1.89s rest follows the bridge phrase.
Q7 TAG/CHORUS ESCALATION: BETA apex identical A3 x4 (no escalation). ALPHA apex A3, B3, A3 (non-monotonic). ALPHA median peak -5.6, -4.9, -4.4 dB = +1.2 dB rise across three statements (weak but correct direction). BETA median peak -7.9, -7.4, -10.1, -8.0 dB (no trend). CHORUS IS ~3 dB QUIETER THAN VERSE throughout.
Q8 CHORD VOCABULARY: 5 distinct roots = Am, G, F, Dm, E. Fully diatonic to A natural minor except the single E major (raised 7th G#) at the bridge - one purposeful secondary dominant. Massively more disciplined than Version A's 22 chords.
Q9 TEMPO/HUMANISATION: melody onsets vs 16th grid at 72 BPM: mean |dev| 52.52ms where random expectation is 52.08ms -> ratio 1.008, INDISTINGUISHABLE FROM RANDOM. Best-fitting tempo by grid search is 144.2 BPM (= 2x 72.1) at 21.36ms. Signed mean -16.68ms (ahead of grid), std 58.81ms. Only 9.5% of durations are exact grid multiples; IOIs within 30ms of a 16th multiple = 40.4% (random expectation 28.8%). IOI histogram clusters at musical values (2.0 beats x19, 0.5 x11, 1.0 x10). CONCLUSION: musical rhythmic values with heavy timing jitter, i.e. strongly humanised, NOT quantised. CAVEAT: detected onsets include synth attack time, so measured jitter is an UPPER BOUND on true MIDI humanisation.
DYNAMIC ARC: regional median-peak spread excluding outro = 6.07 dB. Overall track level envelope is essentially flat until a slow decay after ~170s. ALPHA rises +1.2 dB across three statements. Outro drops to -16.5 dB. So: minimal arc, far weaker than released Version A's 7+ dB section spread.
CAESURAS: the three largest rests are 3.38, 3.38, 3.36s (all exactly 1.01 bars) and each occurs IMMEDIATELY BEFORE a BETA/chorus entry (bars 14, 30, 52). Deliberate and consistent 1-bar rest before the hook - a genuinely good encoded feature.
Corrections
CORRECTION 1 - chorus lift. Two measurements disagree and both are legitimate:
- Section-mean RMS: A +0.70 dB, B +1.42 dB (B better)
- Immediate arrival, loud-frame top-25% over 8s before/after boundary: A +0.72/+2.17/+0.94 (avg +1.28), B +1.53/+0.67/-1.70 (avg +0.17) (A better)
HONEST CONCLUSION: neither version achieves a convincing level lift. Both marginal (under ~2 dB). Level is NOT where B wins. Do not claim B lifts harder on level.
CORRECTION 2 - melodic interval content is NOT a differentiator. Earlier sparse extraction suggested A was leapy and B stepwise. Clean demucs-stem data shows near-identical: A 54% stepwise / 10% leaps, B 57% stepwise / 11% leaps. Median interval 2.0 st both. Do not claim B is more conjunct.
CORRECTION 3 - vocal expression/delivery is essentially IDENTICAL. Vibrato depth A 0.176 st vs B 0.199 st, rate 4.62 vs 4.80 Hz. Note-level dynamic std A 4.66 dB vs B 4.42 dB. Glide 13.9% vs 14.1%. Hard jumps 1.3% both. Suno did NOT add expressive humanity to the delivery. The gains are structural/arrangement/mix, not performance.
CORRECTION 4 - A is the WIDER mix, not B. A side/mid 0.384 vs B 0.252. A widens progressively to 0.521 by the outro. Width is something A does more of, though likely reverb wash rather than intentional placement.
WHERE B GENUINELY WINS (robust across both measurement methods):
- Spectral direction into chorus: A's chorus gets DARKER every time (centroid -143, -69, -36 Hz). B's gets BRIGHTER on 2 of 3 (+6, +247, +224 Hz).
- Vocal/backing balance across choruses: A degrades monotonically (+0.37, -3.25, -5.74 dB) so by the final chorus the hook is ~6 dB UNDER the band. B stays near balanced (+1.15, -3.12, +1.24).
- Harmonic rhythm differentiation: A chorus change-rate 0.83-1.00 = same as its verse (0.92). B drops chorus to 0.62-0.64 vs verse 0.92-1.00. B lets the chorus land and sit.
- Harmonic vocabulary discipline: A 17 unique chords incl. non-functional Amaj7, D7, F#dim, Adim. B 13, diatonic F major core (Dm7 13x, F 11x, Bbmaj7 9x).
- Hook phrase-ending target: ALL SEVEN of B's hook occurrences end on C4. A's end on E4, A3, F4, C4, E4, C4, D4 (no target).
- Bridge as textural event: B bridge perc 52.9% and note density 0.64/beat (30% sparser than verse), medDur 0.66s, 58% notes >0.5s. A bridge perc 27.2% = same as Ch2, density 0.97/beat = same as verse. A has no real bridge.
- Range expansion into chorus: A verse 12.1st -> chorus 13.5st (+1.4). B verse 14.0 -> chorus 19.3 (+5.3).
- Groove: A straight (swing 0.527), vocal locked to backing (+1.2ms), uniform local ibiCV 0.0118-0.0182. B swung (0.561), vocal PUSHES ahead of backing (-13.6ms), section-varying ibiCV 0.0067-0.0448.
- Intro headroom: B intro -28.91 dB and 87.4% of energy in 80-250 Hz (filtered, small) with 1115ms gap before vocal. A intro -24.87 dB, 0ms gap.
SHARED FAILURES (neither version solved):
- Hook motif contour self-similarity near zero in BOTH: A mean pairwise r -0.080, B +0.079. Contour spread 3.60 st vs 3.46 st. Neither sings the hook the same way twice.
- Both deflate in the final chorus: A Ch3 -21.17 vs Ch2 -18.24. B Ch3 -20.18 vs Ch2 -16.55.
- Neither develops the upper spectrum. No band above 500 Hz ever exceeds 3% of energy in ANY section of EITHER version. Both dark/low-mid dominated.
- Both kick-led with almost no hat/cymbal (2%). No real kit.
- Identical macro dynamics: both -14.0 LUFS, both LRA 7.1 LU.
- Both have long final sections (A 50s, B 44s).
=== SONG 2 "Pass of the Shuttle" === A = 243.52s (4:04), B = 239.56s (4:00). Both -13.4 LUFS. A LRA 6.5 LU, B LRA 5.8 LU (B is LESS dynamic on macro measure). A crest 14.50 dB, B 15.56 dB. A RMS std 7.09, B 8.41 (B more short-term dynamic). Confirmed different audio (corr 0.0009). TEMPO: A 71.78 BPM, B 69.84 BPM. B SLOWED ~2.7%. (Song 1: both 69.84.) KEY: both A minor (A 0.887, B 0.856). B did NOT shift centre this time. STEREO: A side/mid 0.4065, B 0.3706. A wider again (same as song 1). SUB-BASS: A 32.6% below 80Hz, B 36.1%. Both far heavier than song 1 (13.3/19.7%).
CORRECTION 5 (song 2) - vocal presence in chorus. Note-level "rest%" said A's chorus was 83-90% rest, implying the vocal nearly vanishes. WRONG - that is a masking artifact. A's chorus backing jumps +9.35 dB, which defeats pyin pitch tracking. Robust stem-energy gate measure instead:
- A: verse 59.0% vocal-present, chorus 58.3% -> -0.7 pts (essentially FLAT, no differentiation)
- B: verse 63.2%, chorus 51.4% -> -11.8 pts (B's chorus is genuinely more instrumental)
Do NOT claim A's chorus vocal disappears. The honest claim is A never differentiates vocal presence at all.
SONG 2 - WHERE A BEATS B (reversal from song 1):
- Verse->chorus level lift: A +3.81 dB vs B +1.96 dB
- Backing lift into Ch1: A +9.35 dB (V1 -27.85 -> Ch1 -18.50) vs B +4.75 dB. A built more verse headroom.
- V3->Ch3 arrival: A +5.47 dB level / +9.24 dB backing / +33.9 pts sub. B only +1.34 / +0.35 / +3.0.
- A's climax is its FINAL chorus (Ch3 -12.94 dB, loudest section). A solved the ending problem that song 1's A failed.
- A's three choruses share a consistent Am-C(maj7)-Fmaj7-C(sus) progression. B's Ch1 differs from Ch2/Ch3.
- A's bridge is genuinely sparse: onset 1.20/s vs 2.46 surrounding (-51%), centroid 1369 Hz (darkest), width 0.295 (narrowest). Supports the tender "bird in the corner" lyric. B's bridge is its BUSIEST section (3.28/s, 47.3% percussive) which fights the lyric.
- A's tag contour is more consistent than B's: A r=+0.624 vs B r=+0.195 (B spread 2.85 st).
SONG 2 - WHERE B BEATS A:
- HOOK consistency: B r=+0.747 (spread 0.87 st) vs A r=+0.480 (spread 1.02 st). Both far better than song 1 (~0). This song HAS a hook in both versions.
- Hook peak note: B peaks A4 on all three occurrences; A peaks E4/F4/F4. B's hook sits a fourth higher with a consistent high target. Corroborated by ceiling measure (A verse E4 -> chorus E4 = -0.7 st; B verse G4 -> chorus A4 = +2.1 st).
- Melodic range: B global C3-C5 (24 st) vs A E3-A4 (17 st). B chorus range expands +1.5 st; A's CONTRACTS -3.1 st.
- Note count 295 vs 249. B centroid 1644 Hz vs A 1502 Hz. B rolloff85 3469 vs 3057.
- B's transitions brighten more (avg ~+840 Hz vs A ~+568 Hz).
- B's tag escalates deliberately: Tag3 is B's loudest section (-13.87), most vocal-present (73.9%), median melody G4 vs Tag1's C4. B uses hook for memorability and tag for escalation.
- B intro>V1 gap 1579 ms vs A's 464 ms.
- B's intro breathes: ibiCV 0.0418, only 5.0% onbeat-anchored. A's intro is a click track: ibiCV 0.0077, 72.4% onbeat.
SONG 2 - SHARED FAILURES (both versions):
- Chorus harmonic rhythm NOT slowed: change-rate 1.00 in verse, chorus AND tag for BOTH. Song 1's key lesson unlearned by both.
- VOCAL BURIED IN EVERY CHORUS, worsening toward the end. A: Ch1 -1.31, Ch2 -1.37, Ch3 -3.09 dB under backing. B: Ch1 -1.39, Ch2 -2.63, Ch3 -3.56. Outro A -3.99, B -4.14. A buried in 9 of 12 sections, B in 6 of 12.
- V1->Ch1 vocal balance swing: A -7.36 dB, B -8.34 dB. Vocal drops 7-8 dB relative to band entering the chorus.
- Nothing above 500 Hz: only ONE section in either version (B's Ch1) has the 500-1000 Hz band above 3% of energy. Same as song 1.
- Both extremely sub-heavy at the end (A Ch3 51.3%, Tag3 42.2%; B Tag3 56.9%, Outro 53.5% below 80 Hz).
- Long outros: A 40s, B 42s.
- A's intro is 24s (10% of song) before first word; B's 14s.
CROSS-SONG SYSTEMIC PATTERNS (both songs, both versions): 1. Vocal buried in choruses, progressively worse toward the end. Appears in 3 of 4 versions analysed (all but song 2's B tags). 2. No energy above 500 Hz anywhere. 4 of 4 versions. 3. Chorus harmonic rhythm never slowed except song 1's B. 3 of 4 versions. 4. A is always the wider mix (song 1: 0.384 vs 0.252; song 2: 0.4065 vs 0.3706). 5. B always brighter than A (song 1: 1896 vs 1802 Hz; song 2: 1644 vs 1502 Hz). 6. B always makes the bridge its most percussive section (song 1: 52.9%; song 2: 47.3%). 7. Long outros in all four versions (40-50s).
=== SONG 4: GUIDE V2 "why does Suno treat it as ambience" === guide-track-v2.mp3: 138.0s (2:18), -17.2 LUFS, LRA 24.3 LU (v1 was 6.3), 128 kbps MP3, L/R corr 0.963. Melody now reaches E4/A4/C5 (v1 topped at D4). 31.9% of frames below -40dB. STARTS WITH 12.45s OF SILENCE.
CORRECTION 6 - MY SEPARATION HYPOTHESIS WAS WRONG. I predicted a Demucs-class separator would route the choir-ooh melody to the ACCOMPANIMENT stem, explaining why Suno ignores it as a vocal. Tested it directly:
- guide v2: 88.92% of energy routed to the VOCALS stem, 11.08% to accompaniment.
- Real Suno mix with an actual lead vocal, same separator: 41.53% vocals / 58.47% accompaniment.
So the guide is MORE vocal-classified than a real mix. Separation is NOT the failure point. Do not claim it is.
REVISED (and better) DIAGNOSIS: Suno is not misclassifying the guide - it is reproducing it faithfully. The guide's acoustic signature is precisely that of BACKGROUND HUMMING, so Suno delivers background humming. Measured against a real lead vocal stem (sounding frames only, >-35dB):
- 2-5kHz presence band: guide 0.126% vs real vocal 1.329% (10.5x deficit)
- 1-2kHz: 0.659% vs 4.115% (6.2x)
- above 5kHz: 0.009% vs 1.852% (206x)
- below 500Hz: 95.94% vs 73.45%
- spectral centroid: 917 Hz vs 2650 Hz (2.9x)
- MFCC std (timbral variety): 9.73 vs 18.42 (1.9x)
- MFCC delta (articulation rate): 1.112 vs 2.202 (2.0x)
- spectral flux: 0.0162 vs 0.0479 (3.0x)
- onsets/sec: 0.92 vs 1.91 (2.1x)
- noise fraction (breath/consonants): 0.16% vs 1.90% (12x)
- spectral flatness: 0.00030 vs 0.01616 (54x)
- zero-crossing rate: 0.0321 vs 0.1059 (3.3x)
- vibrato depth: 0.023 st vs 0.171 st (7.4x)
- voiced fraction: 91.6% vs 81.0% (guide is MORE continuously pitched than a human - no unvoiced consonants)
- longest continuous pitched run: 7.15s vs 6.41s
- spectral envelope movement (formant articulation proxy): std 196 Hz vs 1467 Hz = 7.49x LESS movement in the guide. This is the static "ooh" vowel quantified.
- attack time: 34.7ms (v2) vs 46.2ms real - not a differentiator.
ALSO: guide v2 is LESS vocal-like than v1 on several timbral metrics despite being more musically correct (mfcc_std 9.73 vs 12.76, flux 0.0162 vs 0.0297, noise 0.16% vs 2.59%, presence 0.126% vs 0.340%). v2 improved the music and regressed the timbre.
ALSO: the three layers are timbrally indistinguishable to a separator - the vocals stem is 96.4% in 160-500Hz and the accompaniment stem 81.7% in the same band. Nothing in the file marks one layer as "the lead".
SUNO WORKFLOW RESEARCH (subagent, sourced):
- OFFICIAL: melody preservation is architecturally restricted to Cover and Studio Stem Cover. Audio Influence in Create mode is never documented as a melody-transfer tool.
- Studio Stem Cover (Premier only) explicitly "retains the melody and rhythm of the original stem" - closest match to this use case.
- Cover (Pro+) keeps melody but native UI requires the song be in your own library.
- COMMUNITY CONSENSUS (roo.beehiiv.com May 2026, jackrighteous.com May 2026): slider is 0-100, ~50 is the midpoint and is already "loose inspiration". 65-85 recommended to preserve an uploaded idea. User's 32 is well below the loose-inspiration midpoint.
- Community: keep Weirdness under 25; upload 15-45s (over 60s introduces structural confusion); WAV 44.1/16 not MP3; normalise to -1 dBFS; sparse/solo beats dense mixes for pitch extraction.
- OFFICIAL (Voices article): acapella recordings give cleanest results; if background music is present Suno runs stem extraction first.
- NO EVIDENCE FOUND either way on MIDI-rendered vs real recordings, or mono vs stereo.
Plan Overview
THREAD: iterative A/B analysis to train a local AI composer. Song 1 "The Rinse Light" A vs B - COMPLETE (report cms6hm4cs04ba06adrjilhjzu) Song 2 "Pass of the Shuttle" A vs B - COMPLETE (report cms6jo1yo07mp06adrhessn6h). Recommended a vocal-phrasing guide as next test. Song 3 (CURRENT): user built the vocal-phrasing guide and ran it through Suno. Now wants the GUIDE TRACK ITSELF analysed - what does it encode, not what Suno did with it.
Guide: MIDI-rendered vocal phrasing guide, dark folk, A minor, 72 BPM, "Pass of the Shuttle". Three named instruments: vocal_marker, muted_pulse, room_hum. Fed to Suno as Audio Influence. Reported outcome: guide DID transmit dynamic arc, bridge dropout, final-chorus arrival. Did NOT transmit hook apex, wide chorus melody, chorus harmonic rhythm differentiation, vocal-above-band balance. Goal: determine for each failure whether the guide encoded the wrong thing (authoring error) or Suno ignored a correct encoding (tool limitation).
KEY CONSTRAINT: user supplied MP3 render, not MIDI. Velocities cannot be read directly - only amplitude/level in dB as a proxy. Must state this clearly and offer to parse the MIDI if they have it. Note: user says FOUR chorus statements; my earlier Pass of the Shuttle analysis found three. Verify against the audio. MIDI render advantage: clean synth timbres mean pitch tracking and instrument separation should be far more accurate than on real productions.
Plan Tasks
SONG 1 "The Rinse Light" A vs B - COMPLETE (report cms6hm4cs04ba06adrjilhjzu) SONG 2 "Pass of the Shuttle" A vs B - COMPLETE (report cms6jo1yo07mp06adrhessn6h) SONG 3 GUIDE v1 diagnostic - COMPLETE (report cms6ktlmv08zo06ade8r12isc) SONG 4 GUIDE v2 "why read as ambience" - COMPLETE (report cms6mk1fe00hu06adi92nkvwb)
- [x] Analyse guide v2: 138s, -17.2 LUFS, LRA 24.3, melody now reaches A4/C5, 12.45s leading silence
- [x] Build vocal-likeness metric suite, run on v2 + v1 + real vocal stem + real full mix
- [x] Test separation hypothesis - REFUTED (88.92% of guide routed to vocals stem vs 41.53% for a real mix)
- [x] Measure formant movement via cepstral liftering - 196 Hz sd vs 1467 Hz = 7.5x deficit
- [x] Research Suno Audio Influence via subagent (sourced, labelled official/consensus/anecdote)
- [x] Answer all 5 questions with construction recommendations + pre-flight test thresholds
- [x] Publish figure and report
KEY CONCLUSION: Suno is not misclassifying - it is faithfully reproducing a signal whose acoustic signature IS background humming. Deficits vs a real lead vocal on 13 metrics, 1.9x to 206x. Plus two workflow errors: Create-mode Audio Influence is not the melody-transfer path (Cover/Stem Cover are), and 32 is below the ~50 "loose inspiration" midpoint. TOP RECOMMENDATION: sing/hum the guide into a phone. Fixes all 16 metrics at once.