All guides

Music 8 min August 2026

AI Music Generation Prompting

Free-text composer briefs vs. tag-based prompting, why vocals show up uninvited by default, and the single biggest structural fix — scoring a film in sections instead of one flat bed.

Japesh Singhal

Craft & production

Download skill

AI music generation splits into two genuinely different prompting styles depending on the engine — write it like a brief to a composer, or write it like a list of tags — and getting that backwards is the most common reason a cue comes back wrong. The bigger structural lesson, though, has nothing to do with either style: score a film in sections that match its actual emotional shape, not as one long track laid flat underneath everything.

01

Two genuinely different prompting styles

StyleWrite it likeBest for
Free-text proseA brief to a composer — genre, instrumentation, tempo feel, mood, structure, production character, in full sentencesEngines built for natural-language prompts, generally the strongest for nuanced, specific direction
Tag-basedA comma-separated list of short descriptors, not sentencesEngines built around tags — instrument-specific tags steer the arrangement more reliably here than adjectives do
Do this

"A sparse, patient piano cue. Solo felt piano, close-mic'd with audible key noise, over a low sustained cello drone. Slow — around 68 BPM, rubato, no click. Melancholy but not sentimental; restrained, leaving space. Analogue tape warmth, light room reverb, no percussion, no synths."

This is the free-text style — a real composer's brief, not a keyword list. Say what you don't want in prose too, on any engine that has no negative-prompt field.

Do this

["ambient", "cinematic", "solo felt piano", "sustained cello drone", "melancholy", "slow tempo", "analogue tape warmth", "no percussion"]

The same brief, restyled as tags for a tag-based engine. Roughly 6-12 well-chosen tags consistently outperforms a longer list — past about a dozen, later tags visibly stop landing.

If an engine exposes a numeric BPM parameter, it's the one to reach for when a cue must lock to a known tempo — coarse energy/tempo dials (low/medium/high, slow/medium/fast) exist more widely, but an exact number is rarer and worth using whenever beat-matching across cuts actually matters.

02

Vocals are on by default — say so explicitly, not just in the prompt

Most music-generation engines default to allowing sung vocals, and will readily invent sung words over a cue that was meant as pure underscore. Under dialogue, two voices competing for the same frequency band and the same attention is a real, audible problem.

Set the instrumental toggle, don't rely on the prompt text alone
"Instrumental, no vocals" written in the prompt is a suggestion. If the engine has a dedicated instrumental toggle, set it explicitly on every cue that plays under speech — don't rely on prose alone to suppress vocals.
03

Structured, per-section briefs — a hint, not a contract

For a cue with real internal structure (an intro, a build, a peak, a fall-away), writing it out as timestamped sections in the prompt does steer the arrangement:

A structured brief, written in prose

0:00-0:12 Intro — solo instrument, sparse, no percussion. 0:12-0:35 Build — a second instrument enters underneath, the figure tightens. 0:35-0:50 Peak — a soft rhythm section enters, low strings swell. 0:50-1:00 Fall away — back to the solo instrument, last note left to ring out.

Treat this as a steer, not an exact contract
Writing section timings in prose shapes the arrangement's direction, but most engines won't hit those timestamps to the frame. If a score genuinely needs frame-accurate section boundaries, that's what the next section's per-cue approach is for.
04

Score a film in sections, not one flat bed

The default instinct — generate one long track and lay it under the whole piece — is the wrong one, and it's also the more expensive one.

Generate a separate cue per emotional section instead
It actually hits your structure: a single generated track has its own arbitrary build and drop that won't coincide with your edit, while several shorter cues, each briefed to its own scene, land exactly where the scene lands. It's cheaper on a per-second basis in most pricing models. It's more fixable — if one section is wrong, you regenerate that section, not the whole score. And it keeps dialogue clear — the cue under dialogue can be sparse and strictly instrumental, while the cue over a montage can carry vocals and a full rhythm section. One flat bed forces a single compromise across both.

Brief every cue in a score with the same production vocabulary — same instrumentation, same room character, same processing language — and vary only energy and arrangement between them. That's what makes several separate generations sound like one continuous score instead of three unrelated tracks. Then cross-fade at the cut, or land the change exactly on a picture cut, which hides the seam almost completely.

05

The mixing rule

ContextMusic level under dialogue
Normal narrative / drama12-15 dB below dialogue
Comedy15-20 dB below dialogue — comic timing depends on hearing every consonant and every pause
  • Measure against dialogue peaks, not the loudest music moment — the voice is the reference point.
  • Duck under lines, don't set-and-forget. Pull the cue down further beneath each line and bring it back up in the gaps between them.
  • Brief the arrangement to actually leave room. A cue with a gap in its midrange — a low drone plus high sparkle, nothing in the frequency band a voice actually occupies — sits comfortably under dialogue at a much higher level than a cue that's dense across that exact same band. Say so in the prompt: "sparse midrange, leave room for a voice."
  • When in doubt, quieter. Nobody has ever complained that they could hear the dialogue too clearly.
  • Score the ad's genre, don't generate a music-video bed. Name the cue's job per section. If the client wants music clear from frame 0 under VO, re-lock mix policy (parallel bed vs dialogue-led) instead of silently relaxing the gap rule.
  • Duck with look-ahead. Key the sidechain off dialogue shifted ~150 ms earlier. On a hard cut into a line, a ducker that only reacts after the key arrives lets music peek over the first syllable. When a QC gate fails on a boundary, fix the mix — never relax the threshold.
  • Open-window check. Verify music level in the first ~2 seconds matches the locked policy; trim quiet intros that start after picture.

music covers ElevenLabs' own music prompting specifics in more depth; this page is the cross-engine comparison and the structural scoring technique.