Veo 3.1 is built for one thing above everything else: a single generation call that comes back with picture and audio already locked together — ambient sound, effects, and dialogue arrive synced to the motion, not scored onto it afterward. That's also exactly where its limits show up. This guide covers the prompt formula that gets clean, controllable shots out of it, how to write dialogue it can actually perform, and — the part most guides skip — where its native audio breaks down and what to do instead.
Veo's prompt formula and camera control
Describe the shot the way you'd brief a camera operator: subject, what it's doing, the scene around it, how the camera is moving — in sentences, not tags.
A weathered dock foreman in a soaked oilskin jacket stands at the end of a rain-lashed pier, water dripping off the brim of his cap. Camera slowly pushes in from a wide shot as he turns to face it. Grey overcast light, heavy rain, distant foghorn.
Subject, action, scene, camera, and light are all present as plain description — nothing here is a keyword tag.
Holding a character or product across shots
Give Veo up to three reference images and it will carry a face, an outfit, or a product's exact look into the generated shot instead of reinventing it each time. This is the difference between a character who looks like the same person scene to scene and one who quietly drifts.
First-frame-only vs. locking both ends
Veo can take a starting frame, an end frame, or both, and generate the motion between them. Locking both trades away some of the motion's life for certainty about where the shot lands — useful when the next cut has to match a precise end state, or a title card has to land on an exact frame. But for a shot whose entire job is the motion — fire catching, a hero action, anything where the movement itself is the point — first-frame-only and letting the model originate the motion from a described action reads as more alive. Locking both ends on a motion-driven shot is a common way to end up with something that looks technically correct and slightly frozen.
Built from official sources
- Ekly internal testing across multiple productions; general Veo prompt-formula guidance consistent with Google's own Veo 3/3.1 documentation.
Writing dialogue for Veo's native audio
Veo generates ambient sound, effects, and dialogue together with the picture in one pass — the audio is "native," not scored on afterward. That's powerful when it works and unforgiving when a line comes back wrong, because fixing it means re-rolling the whole shot.
The syntax: no quotes
Write dialogue as A woman says: We have to leave now. — a colon, no quotation marks. This is the opposite convention from some other native-audio video models (Seedance, for one, wants quotes), and getting it backwards on Veo tends to invite a burned-in subtitle rendered into the shot instead of clean spoken audio.
Medium shot of a woman in a rain-soaked coat, urgency in her eyes. She turns to the camera. A woman says: We have to leave now.
Colon, no quotes. Name the speaker by description if there's more than one person in frame — an unattributed line can land on whichever face the model finds more salient.
A woman says "We have to leave now."
Quotation marks on Veo specifically invite a burned-in subtitle instead of clean spoken audio.
Keep the line sayable in the clip's duration
Veo generates in fixed durations — 4, 6, or 8 seconds. A line that takes longer than that to say naturally gets a sped-up, unnatural read, or audio that outruns the picture. Write to the duration, not around it.
Built from official sources
- Ekly internal testing.
Veo's non-English dialogue limits (and what to use instead)
English is the one language where Veo's native dialogue is a safe default. Everything else needs a plan.
A specific, separate problem: script can trip the content filter
Beyond the pronunciation question, a quoted line written in a non-Latin script (Devanagari, for Hindi) combined with heavy performance-direction language in the same prompt has been observed to trip Veo's content-safety filter and block the generation outright — even on entirely ordinary, non-problematic dialogue. The fix that resolved it: transliterate the line into the Latin alphabet (romanized Hindi, for example) and tone down the adjectival performance description around it. Re-sending the same line that way, with no other change, cleared the filter.
What to actually do for a non-English line
Don't reach for a different video model's native voice as the fix — none of them have a confirmed non-English track record either. The reliable move is to decouple the audio: generate the line with a dedicated speech model (where you can actually iterate pronunciation and delivery cheaply), then lip-sync it onto the picture separately. video models compared walks through exactly which approach fits which situation, model by model.
Built from official sources
- Third-party reviewer reports on Veo's non-English dialogue quality (independent testing, not Google's own documentation); Ekly internal testing on the Devanagari/content-filter interaction.
Common Veo mistakes
Text on screen renders wrong
Symptom — You asked for a title card, lower-third, or UI text and got garbled letters
- Don't rely on Veo for on-screen text you actually need to read correctly — it renders text unreliably and you can't fix it without a full re-render. Composite real text over the shot in the edit instead.
Generation rejected or wrong length
Symptom — Requested a duration Veo doesn't support
- Veo generates in 4, 6, or 8 second increments only — not an arbitrary length. Plan your cut to those durations rather than fighting the model for something in between.