AI music generation splits into two genuinely different prompting styles depending on the engine — write it like a brief to a composer, or write it like a list of tags — and getting that backwards is the most common reason a cue comes back wrong. The bigger structural lesson, though, has nothing to do with either style: score a film in sections that match its actual emotional shape, not as one long track laid flat underneath everything.
Two genuinely different prompting styles
"A sparse, patient piano cue. Solo felt piano, close-mic'd with audible key noise, over a low sustained cello drone. Slow — around 68 BPM, rubato, no click. Melancholy but not sentimental; restrained, leaving space. Analogue tape warmth, light room reverb, no percussion, no synths."
This is the free-text style — a real composer's brief, not a keyword list. Say what you don't want in prose too, on any engine that has no negative-prompt field.
["ambient", "cinematic", "solo felt piano", "sustained cello drone", "melancholy", "slow tempo", "analogue tape warmth", "no percussion"]
The same brief, restyled as tags for a tag-based engine. Roughly 6-12 well-chosen tags consistently outperforms a longer list — past about a dozen, later tags visibly stop landing.
If an engine exposes a numeric BPM parameter, it's the one to reach for when a cue must lock to a known tempo — coarse energy/tempo dials (low/medium/high, slow/medium/fast) exist more widely, but an exact number is rarer and worth using whenever beat-matching across cuts actually matters.
Vocals are on by default — say so explicitly, not just in the prompt
Most music-generation engines default to allowing sung vocals, and will readily invent sung words over a cue that was meant as pure underscore. Under dialogue, two voices competing for the same frequency band and the same attention is a real, audible problem.
Structured, per-section briefs — a hint, not a contract
For a cue with real internal structure (an intro, a build, a peak, a fall-away), writing it out as timestamped sections in the prompt does steer the arrangement:
A structured brief, written in prose
0:00-0:12 Intro — solo instrument, sparse, no percussion. 0:12-0:35 Build — a second instrument enters underneath, the figure tightens. 0:35-0:50 Peak — a soft rhythm section enters, low strings swell. 0:50-1:00 Fall away — back to the solo instrument, last note left to ring out.
Score a film in sections, not one flat bed
The default instinct — generate one long track and lay it under the whole piece — is the wrong one, and it's also the more expensive one.
Brief every cue in a score with the same production vocabulary — same instrumentation, same room character, same processing language — and vary only energy and arrangement between them. That's what makes several separate generations sound like one continuous score instead of three unrelated tracks. Then cross-fade at the cut, or land the change exactly on a picture cut, which hides the seam almost completely.
The mixing rule
- Measure against dialogue peaks, not the loudest music moment — the voice is the reference point.
- Duck under lines, don't set-and-forget. Pull the cue down further beneath each line and bring it back up in the gaps between them.
- Brief the arrangement to actually leave room. A cue with a gap in its midrange — a low drone plus high sparkle, nothing in the frequency band a voice actually occupies — sits comfortably under dialogue at a much higher level than a cue that's dense across that exact same band. Say so in the prompt: "sparse midrange, leave room for a voice."
- When in doubt, quieter. Nobody has ever complained that they could hear the dialogue too clearly.
- Score the ad's genre, don't generate a music-video bed. Name the cue's job per section. If the client wants music clear from frame 0 under VO, re-lock mix policy (parallel bed vs dialogue-led) instead of silently relaxing the gap rule.
- Duck with look-ahead. Key the sidechain off dialogue shifted ~150 ms earlier. On a hard cut into a line, a ducker that only reacts after the key arrives lets music peek over the first syllable. When a QC gate fails on a boundary, fix the mix — never relax the threshold.
- Open-window check. Verify music level in the first ~2 seconds matches the locked policy; trim quiet intros that start after picture.
music covers ElevenLabs' own music prompting specifics in more depth; this page is the cross-engine comparison and the structural scoring technique.