# The Complete Guide to Prompting Veo 3.1 — a prompting skill for Veo 3.1
Source: ekly.ai/guides/veo-prompting — free to use with any AI assistant.

## Hard rules

1. Describe shots in plain sentences: subject, action, scene, camera, light — never a keyword tag list.
2. Use up to three reference images when a character or product must hold across shots.
3. Lock first-frame-only when the shot's job is the motion. Lock both ends only when the next cut needs an exact end frame.
4. Write dialogue as `Speaker says: line` with a colon and no quotation marks. Quotes invite burned-in subtitles on Veo.
5. Keep every spoken line sayable inside the clip duration — Veo generates only 4, 6, or 8 seconds.
6. Treat English as the safe default for Veo's native dialogue. Outside English, plan to decouple audio.
7. If a non-English SPEAK beat is blocked with Devanagari (or another non-Latin script) plus heavy performance language: romanize the line and soften performance adjectives before assuming a real content violation.
8. Do not switch video models hoping their native voice is better for non-English — none have a confirmed non-English track record. Generate speech separately, then lip-sync.
9. Never rely on Veo for readable on-screen text — composite real text in post.
10. Request only durations of 4, 6, or 8 seconds.

## Prompt formula

Brief a camera operator: who is in frame, what they are doing, where they are, how the camera moves, what the light is. Example shape:

> A weathered dock foreman in a soaked oilskin jacket stands at the end of a rain-lashed pier, water dripping off the brim of his cap. Camera slowly pushes in from a wide shot as he turns to face it. Grey overcast light, heavy rain, distant foghorn.

## Dialogue syntax

- Good: `A woman says: We have to leave now.`
- Bad: `A woman says "We have to leave now."`

Name the speaker by description when more than one person is in frame.

## Non-English dialogue (reported pattern)

Independent reviewers report Veo's dialogue quality drops outside English — inconsistent pronunciation, flatter delivery, unstable accent. Google publishes no per-language quality statement; treat the pattern as strong enough to plan around, not as an official spec.

For non-English lines: generate with a dedicated speech model, then lip-sync onto picture. See ekly.ai/guides/multilingual-dialogue#video-models-compared.

## Common mistakes

- On-screen text garbles → composite in edit.
- Arbitrary duration → only 4 / 6 / 8 seconds.
