All guides

Video models 14 min August 2026

The Complete Guide to Prompting Veo 3.1

Google's cinematic video model generates picture and synced audio together — here's how to direct the shot, write the dialogue, and know when to route audio around it instead of through it.

Japesh Singhal

Craft & production

Download skill

Veo 3.1 is built for one thing above everything else: a single generation call that comes back with picture and audio already locked together — ambient sound, effects, and dialogue arrive synced to the motion, not scored onto it afterward. That's also exactly where its limits show up. This guide covers the prompt formula that gets clean, controllable shots out of it, how to write dialogue it can actually perform, and — the part most guides skip — where its native audio breaks down and what to do instead.

Four things to get right
Subject and camera direction in plain, specific sentences (not a keyword list). Reference images when a character or product has to hold across shots. First-frame-only when the shot's whole job is the motion; both ends locked only when hitting an exact end frame matters more. And know before you generate whether the dialogue's language is one Veo's native audio actually handles well.
01

Veo's prompt formula and camera control

Describe the shot the way you'd brief a camera operator: subject, what it's doing, the scene around it, how the camera is moving — in sentences, not tags.

Do this

A weathered dock foreman in a soaked oilskin jacket stands at the end of a rain-lashed pier, water dripping off the brim of his cap. Camera slowly pushes in from a wide shot as he turns to face it. Grey overcast light, heavy rain, distant foghorn.

Subject, action, scene, camera, and light are all present as plain description — nothing here is a keyword tag.

Holding a character or product across shots

Give Veo up to three reference images and it will carry a face, an outfit, or a product's exact look into the generated shot instead of reinventing it each time. This is the difference between a character who looks like the same person scene to scene and one who quietly drifts.

First-frame-only vs. locking both ends

Veo can take a starting frame, an end frame, or both, and generate the motion between them. Locking both trades away some of the motion's life for certainty about where the shot lands — useful when the next cut has to match a precise end state, or a title card has to land on an exact frame. But for a shot whose entire job is the motion — fire catching, a hero action, anything where the movement itself is the point — first-frame-only and letting the model originate the motion from a described action reads as more alive. Locking both ends on a motion-driven shot is a common way to end up with something that looks technically correct and slightly frozen.

Rule of thumb
If the next edit needs this shot to land on a specific frame, lock both ends. If nobody's cutting into anything and the shot just needs to breathe, lock the start and describe the motion in prose.

Built from official sources

  • Ekly internal testing across multiple productions; general Veo prompt-formula guidance consistent with Google's own Veo 3/3.1 documentation.
02

Writing dialogue for Veo's native audio

Veo generates ambient sound, effects, and dialogue together with the picture in one pass — the audio is "native," not scored on afterward. That's powerful when it works and unforgiving when a line comes back wrong, because fixing it means re-rolling the whole shot.

The syntax: no quotes

Write dialogue as A woman says: We have to leave now. — a colon, no quotation marks. This is the opposite convention from some other native-audio video models (Seedance, for one, wants quotes), and getting it backwards on Veo tends to invite a burned-in subtitle rendered into the shot instead of clean spoken audio.

Do this

Medium shot of a woman in a rain-soaked coat, urgency in her eyes. She turns to the camera. A woman says: We have to leave now.

Colon, no quotes. Name the speaker by description if there's more than one person in frame — an unattributed line can land on whichever face the model finds more salient.

Not this

A woman says "We have to leave now."

Quotation marks on Veo specifically invite a burned-in subtitle instead of clean spoken audio.

Keep the line sayable in the clip's duration

Veo generates in fixed durations — 4, 6, or 8 seconds. A line that takes longer than that to say naturally gets a sped-up, unnatural read, or audio that outruns the picture. Write to the duration, not around it.

Built from official sources

  • Ekly internal testing.
03

Veo's non-English dialogue limits (and what to use instead)

English is the one language where Veo's native dialogue is a safe default. Everything else needs a plan.

Reported, not Google's own documentation
Independent reviewers consistently report that Veo's dialogue quality drops outside English — inconsistent pronunciation, flatter emotional delivery, unstable accent and cultural direction. Google publishes no per-language quality statement of its own, so treat this as a strong, consistent pattern across testers rather than an official spec. It's still reliable enough to plan around.

A specific, separate problem: script can trip the content filter

Beyond the pronunciation question, a quoted line written in a non-Latin script (Devanagari, for Hindi) combined with heavy performance-direction language in the same prompt has been observed to trip Veo's content-safety filter and block the generation outright — even on entirely ordinary, non-problematic dialogue. The fix that resolved it: transliterate the line into the Latin alphabet (romanized Hindi, for example) and tone down the adjectival performance description around it. Re-sending the same line that way, with no other change, cleared the filter.

If a non-English SPEAK beat gets blocked
Before assuming the content itself is the problem, try romanizing the script and softening performance-direction language first. This reads as a script/filter interaction, not a real content violation.

What to actually do for a non-English line

Don't reach for a different video model's native voice as the fix — none of them have a confirmed non-English track record either. The reliable move is to decouple the audio: generate the line with a dedicated speech model (where you can actually iterate pronunciation and delivery cheaply), then lip-sync it onto the picture separately. video models compared walks through exactly which approach fits which situation, model by model.

Built from official sources

  • Third-party reviewer reports on Veo's non-English dialogue quality (independent testing, not Google's own documentation); Ekly internal testing on the Devanagari/content-filter interaction.
04

Common Veo mistakes

Text on screen renders wrong

Symptom — You asked for a title card, lower-third, or UI text and got garbled letters

  • Don't rely on Veo for on-screen text you actually need to read correctly — it renders text unreliably and you can't fix it without a full re-render. Composite real text over the shot in the edit instead.

Generation rejected or wrong length

Symptom — Requested a duration Veo doesn't support

  • Veo generates in 4, 6, or 8 second increments only — not an arbitrary length. Plan your cut to those durations rather than fighting the model for something in between.