Thread Context

29 Jul 2026

Thread Context

Key facts and notes for this thread. Updated by agent, survives context compaction.

Numbers & Values

FILES: A = "The Rinse Light -A.mp3" (231.55s / 3:51). B = "The Rinse Light -B.mp3" (208.99s / 3:29). B is 22.6s shorter. Confirmed genuinely different audio (zero-lag corr -0.0007).

TEMPO: both share fundamental 69.84 BPM (autocorrelation strongest peak both files). B preserved A's tempo. Beat salience stronger in B (19171 vs 16896). Beat-interval CV: A 0.0551, B 0.0573. Tempo drift across song: A +1.05 BPM (speeds up), B -2.74 BPM (relaxes).

KEY: A reads C major (0.913) / A minor (0.825). B reads F major (0.920) / D minor (0.673). Tonal centre shifted.

LOUDNESS: both normalised to -14.0 LUFS integrated, both LRA 7.1 LU. Crest factor A 16.52 dB, B 15.22 dB. Dynamic span p5-p95: A 21.66 dB, B 21.97 dB. Macro dynamic range essentially IDENTICAL - B did not expand it.

STEREO: A side/mid 0.384, L/R corr 0.743 (wider). B side/mid 0.252, L/R corr 0.881 (narrower/more centred). A is the wider mix overall.

VERSE-TO-CHORUS LIFT: A +0.70 dB, centroid -345 Hz, sub-bass +7.9 pts, onsets +15%. B +1.42 dB, centroid -78 Hz, sub-bass +16.0 pts, onsets +1%. Both weak on level; A roughly half of B.

SECTION LEVEL ARC (RMS dB rel peak): A: intro -24.87, V1 -22.27, Ch1a -19.55, Ch1b -18.80, V2 -20.08, Pre -18.07, Ch2 -18.24, Bridge -20.81, Ch3+outro -21.17. Peak at Pre/Ch2; final chorus 2.9 dB QUIETER than Ch2. B: intro -28.91, V1 -22.06, Ch1 -17.60, V2 -17.01, Ch2 -16.55, Bridge -17.98, Ch3+outro -20.18. Monotonic build to Ch2 then release.

SUB-BASS % BY SECTION: A: 1.2, 6.2, 7.8, 10.4, 7.7, 4.6, 15.9, 18.7, 22.0 B: 2.1, 4.9, 13.1, 10.0, 31.3, 30.7, 26.1 (B commits ~2x sub in choruses)

STEREO WIDTH BY SECTION: A 0.212 to 0.521 (progressive widening, ends widest). B 0.192 to 0.320.

PERCUSSIVE RATIO: A bridge 17.1%. B bridge 27.5% (B makes bridge a rhythmic event).

ONSET MICRO-TIMING vs 16th grid: A mean abs dev 13.4 ms, std 20.7 ms. B 11.7 ms, std 18.7 ms. B tighter.

SPECTRAL: centroid A 1802 Hz, B 1896 Hz. Rolloff85 A 3954, B 4159. Flatness A 0.00972, B 0.01199.

TRANSCRIPT EMOTION TAGS: A = sad on all 8 segments (uniform). B = neutral/sad/neutral/sad/sad/neutral (differentiates verse vs chorus).

STRUCTURE: A has 9 sections incl. split chorus halves (Ch1a 57-71, Ch1b 71-97) and separate Pre/V2b (124-137). B merged chorus halves into single 33s chorus and merged V2+pre into one 35s verse. Same lyric, same 281 word count.

LYRIC VARIANTS: A "like I forgot to close" vs B "like it forgot to close". A "I lined the coins" vs B "I line the coins". A "the amber word rinse" vs B transcribed "amber wood rinse". A "worse prayers" vs B "worst prayers" (B diction less clear on those two).

=== SONG 3: THE GUIDE TRACK ITSELF (guide-track-upload-safe.mp3) === 196.03s (3:16). -18.6 LUFS, LRA 6.3 LU, mean -22.4 dB, max -1.6 dB. L/R corr 0.956, side/mid 0.180 (near-mono). 72 BPM -> bar 3.3333s, beat 0.8333s, 58.8 bars total.

THREE LAYERS IDENTIFIED (MP3 render, so levels are amplitude not MIDI velocity): 1. vocal_marker = melody. 105 notes, range C3-D4 (14 st), median G3, median dur 1.300s (1.56 beats), 78% >0.5s, 57% >1s. Each note preceded by a ~23-35ms portamento glide 1 semitone below (70% of 66 short mid events are exactly +1 st below the following melody note, median dt -23ms). Median CQT peak -5.86 dB. 2. room_hum = SUSTAINED MOVING BASS LINE, not a static hum. Roots A2 x7, G2 x7, F2 x3, D2 x1, E2 x1. Durations 2 bars (6.6s) or 4 bars (13.35s). Median CQT peak -3.75 dB = LOUDEST layer. 3. muted_pulse = quiet high clicks C#6 (43), B5 (42), G5 (10). Median peak -32.43 dB. IOI median 0.406s (~8th note). MONOPHONIC: 166 of 169 groups are single notes. NO CHORD STABS ANYWHERE.

BALANCE: melody minus bass = -2.11 dB (melody is BELOW the loudest backing element). Melody minus pulse = +26.57 dB. In its own band (126-300Hz) melody is +28.94 dB above bass harmonic bleed, so it is clearly audible - not masked, but not dominant either.

STRUCTURE (derived from melodic repeats + bass harmony, mutually confirming): bars 1-7 (0-22.7s) opening. NO INSTRUMENTAL INTRO - first melody note at t=0.000, downbeat of bar 1. bars 8-13 (24.5-42.0) VERSE 1 = ALPHA motif, 11 notes, over 4-bar A2 pedal bars 14-16 (45.4-52.0) CHORUS 1 = BETA motif "A3 G3 D3" + long G3 (2.87s), over F2(2 bars)->G2(2 bars) bars 17-20 (53.3-65.2) TAG 1 bars 21-24 (66.7-76.1) link bars 25-29 (77.6-95.7) VERSE 2 = ALPHA-2 over 4-bar A2 pedal bars 30-32 (99.0-105.7) CHORUS 2 = BETA-2 over F2->G2 bars 33-36 (107-119) TAG 2 (contains the single D4) bars 37-38 (120.4-125.1) BRIDGE, E2 bass = E major, contains the ONLY G#3 in the track bars 39-41 (127-135.6) bridge continuation bars 42-43 (137.4-143) CHORUS 3 variant bars 44-45 (144.5-149.7) link bars 46-51 (151.7-169.2) VERSE 3 = ALPHA-3 over 4-bar A2 pedal bars 52-54 (172.6-179.1) CHORUS 4 = BETA-4 bars 55-59 (180.5-189) outro The 4-bar A2 pedal aligns exactly with each ALPHA; F2->G2 2-bar changes align exactly with each BETA. Mutual confirmation of the map.

Q1 HOOK APEX: BETA apex = A3 on ALL FOUR statements. Contour self-similarity r = +1.000, spread 0.00 st (pitch sequence A3-G3-D3 literally identical each time). BUT A3 is only the 10th-highest pitch used; the track reaches C4 8x and D4 1x, all OUTSIDE the chorus. So the apex is perfectly consistent but 3 semitones below the track's own ceiling and identical to the verse's apex. ALPHA (verse) apex: A3, B3, A3 - varies by 2 st.

Q2 RANGE: VERSE (ALPHA) span 9, 11, 9 st. CHORUS (BETA) span 7, 7, 7, 7 st. Chorus is NARROWER than verse by ~2-4 st. Guide DOES encode chorus contraction. Global range C3-D4 = 14 st. D4 occurs exactly ONCE (bar 33, in a tag).

Q3 HARMONIC RHYTHM: verse = 4-bar A minor pedal = 0.25 chords/bar. Chorus = F 2 bars then G 2 bars = 0.5 chords/bar. CHORUS IS 2x FASTER THAN VERSE - inverted differentiation. Also the chorus never touches the tonic (F->G plagal, unresolved).

Q4 BALANCE: see above. Vocal 2.11 dB BELOW bass, 26.57 dB above pulse.

Q5 INTRO: ZERO. First melody note at t=0.000s on the downbeat of bar 1.

Q6 BRIDGE (bars 37-38): notes/beat 0.48 vs neighbours 0.38 and 0.42 -> density INCREASES. rest 12% = LOWEST of any section. medDur 1.68s = longest. span 4 st = narrowest. So sustained and narrow but NOT sparse; no dropout encoded. A 1.89s rest follows the bridge phrase.

Q7 TAG/CHORUS ESCALATION: BETA apex identical A3 x4 (no escalation). ALPHA apex A3, B3, A3 (non-monotonic). ALPHA median peak -5.6, -4.9, -4.4 dB = +1.2 dB rise across three statements (weak but correct direction). BETA median peak -7.9, -7.4, -10.1, -8.0 dB (no trend). CHORUS IS ~3 dB QUIETER THAN VERSE throughout.

Q8 CHORD VOCABULARY: 5 distinct roots = Am, G, F, Dm, E. Fully diatonic to A natural minor except the single E major (raised 7th G#) at the bridge - one purposeful secondary dominant. Massively more disciplined than Version A's 22 chords.

Q9 TEMPO/HUMANISATION: melody onsets vs 16th grid at 72 BPM: mean |dev| 52.52ms where random expectation is 52.08ms -> ratio 1.008, INDISTINGUISHABLE FROM RANDOM. Best-fitting tempo by grid search is 144.2 BPM (= 2x 72.1) at 21.36ms. Signed mean -16.68ms (ahead of grid), std 58.81ms. Only 9.5% of durations are exact grid multiples; IOIs within 30ms of a 16th multiple = 40.4% (random expectation 28.8%). IOI histogram clusters at musical values (2.0 beats x19, 0.5 x11, 1.0 x10). CONCLUSION: musical rhythmic values with heavy timing jitter, i.e. strongly humanised, NOT quantised. CAVEAT: detected onsets include synth attack time, so measured jitter is an UPPER BOUND on true MIDI humanisation.

DYNAMIC ARC: regional median-peak spread excluding outro = 6.07 dB. Overall track level envelope is essentially flat until a slow decay after ~170s. ALPHA rises +1.2 dB across three statements. Outro drops to -16.5 dB. So: minimal arc, far weaker than released Version A's 7+ dB section spread.

CAESURAS: the three largest rests are 3.38, 3.38, 3.36s (all exactly 1.01 bars) and each occurs IMMEDIATELY BEFORE a BETA/chorus entry (bars 14, 30, 52). Deliberate and consistent 1-bar rest before the hook - a genuinely good encoded feature.

Corrections

CORRECTION 1 - chorus lift. Two measurements disagree and both are legitimate:

CORRECTION 2 - melodic interval content is NOT a differentiator. Earlier sparse extraction suggested A was leapy and B stepwise. Clean demucs-stem data shows near-identical: A 54% stepwise / 10% leaps, B 57% stepwise / 11% leaps. Median interval 2.0 st both. Do not claim B is more conjunct.

CORRECTION 3 - vocal expression/delivery is essentially IDENTICAL. Vibrato depth A 0.176 st vs B 0.199 st, rate 4.62 vs 4.80 Hz. Note-level dynamic std A 4.66 dB vs B 4.42 dB. Glide 13.9% vs 14.1%. Hard jumps 1.3% both. Suno did NOT add expressive humanity to the delivery. The gains are structural/arrangement/mix, not performance.

CORRECTION 4 - A is the WIDER mix, not B. A side/mid 0.384 vs B 0.252. A widens progressively to 0.521 by the outro. Width is something A does more of, though likely reverb wash rather than intentional placement.

WHERE B GENUINELY WINS (robust across both measurement methods):

SHARED FAILURES (neither version solved):

=== SONG 2 "Pass of the Shuttle" === A = 243.52s (4:04), B = 239.56s (4:00). Both -13.4 LUFS. A LRA 6.5 LU, B LRA 5.8 LU (B is LESS dynamic on macro measure). A crest 14.50 dB, B 15.56 dB. A RMS std 7.09, B 8.41 (B more short-term dynamic). Confirmed different audio (corr 0.0009). TEMPO: A 71.78 BPM, B 69.84 BPM. B SLOWED ~2.7%. (Song 1: both 69.84.) KEY: both A minor (A 0.887, B 0.856). B did NOT shift centre this time. STEREO: A side/mid 0.4065, B 0.3706. A wider again (same as song 1). SUB-BASS: A 32.6% below 80Hz, B 36.1%. Both far heavier than song 1 (13.3/19.7%).

CORRECTION 5 (song 2) - vocal presence in chorus. Note-level "rest%" said A's chorus was 83-90% rest, implying the vocal nearly vanishes. WRONG - that is a masking artifact. A's chorus backing jumps +9.35 dB, which defeats pyin pitch tracking. Robust stem-energy gate measure instead:

SONG 2 - WHERE A BEATS B (reversal from song 1):

SONG 2 - WHERE B BEATS A:

SONG 2 - SHARED FAILURES (both versions):

CROSS-SONG SYSTEMIC PATTERNS (both songs, both versions): 1. Vocal buried in choruses, progressively worse toward the end. Appears in 3 of 4 versions analysed (all but song 2's B tags). 2. No energy above 500 Hz anywhere. 4 of 4 versions. 3. Chorus harmonic rhythm never slowed except song 1's B. 3 of 4 versions. 4. A is always the wider mix (song 1: 0.384 vs 0.252; song 2: 0.4065 vs 0.3706). 5. B always brighter than A (song 1: 1896 vs 1802 Hz; song 2: 1644 vs 1502 Hz). 6. B always makes the bridge its most percussive section (song 1: 52.9%; song 2: 47.3%). 7. Long outros in all four versions (40-50s).

=== SONG 4: GUIDE V2 "why does Suno treat it as ambience" === guide-track-v2.mp3: 138.0s (2:18), -17.2 LUFS, LRA 24.3 LU (v1 was 6.3), 128 kbps MP3, L/R corr 0.963. Melody now reaches E4/A4/C5 (v1 topped at D4). 31.9% of frames below -40dB. STARTS WITH 12.45s OF SILENCE.

CORRECTION 6 - MY SEPARATION HYPOTHESIS WAS WRONG. I predicted a Demucs-class separator would route the choir-ooh melody to the ACCOMPANIMENT stem, explaining why Suno ignores it as a vocal. Tested it directly:

REVISED (and better) DIAGNOSIS: Suno is not misclassifying the guide - it is reproducing it faithfully. The guide's acoustic signature is precisely that of BACKGROUND HUMMING, so Suno delivers background humming. Measured against a real lead vocal stem (sounding frames only, >-35dB):

ALSO: guide v2 is LESS vocal-like than v1 on several timbral metrics despite being more musically correct (mfcc_std 9.73 vs 12.76, flux 0.0162 vs 0.0297, noise 0.16% vs 2.59%, presence 0.126% vs 0.340%). v2 improved the music and regressed the timbre.

ALSO: the three layers are timbrally indistinguishable to a separator - the vocals stem is 96.4% in 160-500Hz and the accompaniment stem 81.7% in the same band. Nothing in the file marks one layer as "the lead".

SUNO WORKFLOW RESEARCH (subagent, sourced):

Plan Overview

THREAD: iterative A/B analysis to train a local AI composer. Song 1 "The Rinse Light" A vs B - COMPLETE (report cms6hm4cs04ba06adrjilhjzu) Song 2 "Pass of the Shuttle" A vs B - COMPLETE (report cms6jo1yo07mp06adrhessn6h). Recommended a vocal-phrasing guide as next test. Song 3 (CURRENT): user built the vocal-phrasing guide and ran it through Suno. Now wants the GUIDE TRACK ITSELF analysed - what does it encode, not what Suno did with it.

Guide: MIDI-rendered vocal phrasing guide, dark folk, A minor, 72 BPM, "Pass of the Shuttle". Three named instruments: vocal_marker, muted_pulse, room_hum. Fed to Suno as Audio Influence. Reported outcome: guide DID transmit dynamic arc, bridge dropout, final-chorus arrival. Did NOT transmit hook apex, wide chorus melody, chorus harmonic rhythm differentiation, vocal-above-band balance. Goal: determine for each failure whether the guide encoded the wrong thing (authoring error) or Suno ignored a correct encoding (tool limitation).

KEY CONSTRAINT: user supplied MP3 render, not MIDI. Velocities cannot be read directly - only amplitude/level in dB as a proxy. Must state this clearly and offer to parse the MIDI if they have it. Note: user says FOUR chorus statements; my earlier Pass of the Shuttle analysis found three. Verify against the audio. MIDI render advantage: clean synth timbres mean pitch tracking and instrument separation should be far more accurate than on real productions.

Plan Tasks

SONG 1 "The Rinse Light" A vs B - COMPLETE (report cms6hm4cs04ba06adrjilhjzu) SONG 2 "Pass of the Shuttle" A vs B - COMPLETE (report cms6jo1yo07mp06adrhessn6h) SONG 3 GUIDE v1 diagnostic - COMPLETE (report cms6ktlmv08zo06ade8r12isc) SONG 4 GUIDE v2 "why read as ambience" - COMPLETE (report cms6mk1fe00hu06adi92nkvwb)

KEY CONCLUSION: Suno is not misclassifying - it is faithfully reproducing a signal whose acoustic signature IS background humming. Deficits vs a real lead vocal on 13 metrics, 1.9x to 206x. Plus two workflow errors: Create-mode Audio Influence is not the melody-transfer path (Cover/Stem Cover are), and 32 is below the ~50 "loose inspiration" midpoint. TOP RECOMMENDATION: sing/hum the guide into a phone. Fixes all 16 metrics at once.

Back to studies