Listening Notes - Suno vs ACE 1.5 (2026-08-17)
Listening Notes - Suno vs ACE 1.5 (2026-08-17)
Track: "The Hand That Stays Out"
Objective Measurements
|--------|------|---------| | Duration | 3:05 (184.6s) | 3:22 (202.0s) | | Integrated loudness | -13.5 LUFS | -14.5 LUFS | | True peak | -0.7 dBTP | -1.2 dBTP | | Loudness range (LRA) | 6.1 LU | 5.3 LU | | Overall RMS | -15.2 dB | -16.1 dB | | Peak | -0.7 dB | -1.3 dB |
Section-by-Section RMS (10s windows)
|------|----------|------------|-------| | 0s | -19.5 | -19.8 | Both start quiet - intro/verse 1 | | 10s | -18.0 | -17.7 | Both rising - verse 1 continues | | 40s | -16.8 | -17.2 | Verse 1 → chorus 1 transition | | 60s | -15.6 | -16.1 | | | 80s | -16.3 | -16.1 | | | 90s | -15.9 | -19.5 | ACE drops - possible verse 2 / quieter section | | 150s | -14.3 | -14.1 | Both in outro range | | 160s | -15.8 | -14.7 | ACE slightly louder (outro continuation) | | 180s | - | -14.8 | ACE only (longer track) | | 190s | - | -18.5 | ACE final fade |
Key Observations
- The Turn at ~110s is clearly visible - RMS drops to -23.1 dB, the lowest point in the track. The bed genuinely drops out.
- Chorus 2 at 130-140s is the loudest section (-14.0 to -13.7 dB) - the B4 climax on "stays out" is real.
- The dynamic architecture matches the plan: flat verses (-15 to -17), Turn drop (-23), C2 peak (-13.7), outro (-12.5 at the very end).
- The -23.1 dB drop at 110s is a 3x RMS cliff from the verse level (~-15.6) - exactly the breathturn the composition plan specified.
ACE 1.5 track:
- More consistent level overall - less dramatic dynamic shifts.
- The track is 17 seconds longer (202 vs 184.6s) - ACE may be rendering at a slightly different tempo or extending sections.
- The outro fades more gradually (170s -15.6 → 190s -18.5).
Gemma4:e2b Listening (verse/clip section)
- Detected: acoustic guitar, clear intimate vocal, narrative/conversational style
- Vocal: front and center, gentle and reflective
- Tempo: moderate to slow
- Mood: warm, nostalgic, reflective
- Arrangement: sparse, focused on guitar + vocal
- Production: warm, natural, clean, intimate
- Genre: Folk / Acoustic Singer-Songwriter
ACE 1.5 (55-85s clip):
- Detected: acoustic guitar, warm conversational vocal
- Similar folk/singer-songwriter characterisation
- Notable: gemma flagged "organic imperfections" and "human warmth" as differentiators from typical AI music
- Described as having "genuine human warmth" that AI models struggle to replicate
Comparison Summary
1. Both tracks are in the same genre - folk/singer-songwriter with acoustic guitar and intimate vocal. The composition translated well to both renderers.
3. ACE is more consistent - less variation between sections, gentler transitions. This could be a strength (consistency) or a weakness (less emotional impact from the breathturn).
4. ACE is longer - 17 seconds longer. Either the tempo is slightly slower or sections are extended. This needs investigation.
5. ACE has no Dakota voice - the vocal is present but generic. Once we train the voice model, this becomes the differentiator.
Next Steps
- Listen to both full tracks to validate the RMS section analysis
- Compare ACE tempo vs the 118 BPM plan (is it rendering at a different speed?)
- When voice training is ready, re-render with Dakota's voice
- Consider ACE as the primary production path for songs where arrangement control matters