Two things became clear after enough failed generations to actually count them: structural fixes hold and textual fixes loop — every problem ever permanently solved was solved by changing the shape of a generation call, never by editing prompt text harder. And a pack of rules that only prevents mistakes produces safe, dead films — continuity is the floor, not the ceiling. These 14 rules are what's left after both lessons were taken seriously.
The hard rules
Platform, then story, before pixels
Two facts come before the idea: can the viewer leave (and how fast), and will they hear it. Score the story before spending any generation budget — a story problem found here costs a sentence; found after generation it costs the whole shot.
e.g. “TikTok/Reels = 0s exit + sound-on; YouTube skippable = 5s exit — decide those before the idea”
Set the genre dials before shot planning
Anything left unset is set by the model's defaults, and the defaults are identical for every genre. That's why unset films all look alike.
e.g. “Comedy ad: comedy dials + one commercial dial at the end — never two full dial sets at once”
Lock the world as pixels
A location appearing in more than one shot must exist as an image before the second shot is generated. Prose cannot lock a place.
e.g. “Empty location plate first ("empty platform, white tile, fluorescent lights") — then every shot references that image”
One speaker, one line, one intention per clip
Two speakers in one generation causes line bleed. Two intentions in eight seconds renders neither.
e.g. “One speaker, one ~3–5s line, mouth closed after — two speakers in one clip bleed lines”
Decide the audio route before generating
Native (the video model voices it — one re-roll per fix), post-dub + lip-sync (generate with the line spoken, replace the voice — iteration costs a fraction of a credit), or audio-driven (the model performs to audio you supply). Choose on cost of iteration, not quality alone.
e.g. “Directed delivery or non-English → post-dub + lip-sync (cents/line), not native (full re-roll)”
Put every constraint where it belongs
Pixels, then parameter, then call shape, then a verification check, then prose — in that order of preference. Cap prose constraints at 8; whatever stays in prose is what you'll loop on.
e.g. “Prop count → plate + verification gate; spoken language → gate — keep prose under 8 constraints”
Never name the artefact you're avoiding
Most video models have no negative prompt, so "no second pair of glasses" is a positive cue for glasses. Rewrite every forbid as an observable positive state.
e.g. “"sunglasses resting on top of his head, eyes fully visible" — not "no second pair"”
A regeneration is a delta, not a rewrite
Change one variable. If the fix can't be expressed as one change, the shot needs splitting into two.
e.g. “sunglasses.placement → resting on head — change that field only; print the diff”
One beat, one review
Never generate the next shot while one awaits review. Batching produces bundled multi-defect feedback, which forces multi-axis fixes — which is what causes regressions.
e.g. “Approve shot N before generating N+1 — language fix and sunglasses fix are separate deltas”
Look before you claim
Never say a clip is correct unless you actually looked at it. If you couldn't look, say that instead.
e.g. “If you didn't watch the take, say "unchecked" — never mark KEEP from the prompt alone”
Two failures on one axis means the tool is wrong
Stop editing the prompt. Switch model or pipeline stage instead — and revisit an earlier tool preference once evidence contradicts it.
e.g. “Language failed twice on this model → switch model or stage; stop rewriting the same prompt”
When a loop is reported, deliver a check, not a paragraph
Don't add a rule to a process unless you can name the specific check that enforces it.
e.g. “Gate: "prop count matches the plate" — not another paragraph of rules”
Direction review before continuity review
Run the silent test (picture only) and the muted test (performance only) before checking props and wardrobe.
e.g. “Silent test (story readable?) → muted test (who wants what?) → then props and wardrobe”
Finish is a fixed pass
Identical every time. Improvised last-minute settings are how films ship with clipped dialogue or the wrong pixel format.
e.g. “Cadence → grade → texture → one encode per delivery spec → QC — same order every film”
Cost-control gates before spend
Verbatim canon lines (not concept bullets); workwear brand mark baked into character sheets; name glossary for ambiguous labels; mix/EDL-only asks never regenerate KEEP picture; SPEAK ends after two syllable fails — mute and dub.
e.g. “Brief says "phone thought" only → RECOMMENDATION line + approval flag, not a silent "locked" VO”
Ship gates before platform-complete
Frame-snap the EDL; beat inventory matches last KEEP; voice map present; open-music readable from t=0 when required; cutdowns, burned captions, and per-surface loudness exist before CLOSE. A hero master alone is not enough.
e.g. “30s 16:9 KEEP'd → still need planned 15s/6s and 9:16/4:5 cutdowns + caption pass”
The pipeline
The default order of operations, start to finish. Skipping a stage doesn't save time — it moves the cost of finding the problem later, where it's more expensive:
The default pipeline
Brief → Platform (exit pressure, sound state, hero length) → Story (shape, four movements, score it) → Platform again (full surface list, hero ratio) → Genre (set the dials) → Dialogue (write it, then cut the first sentence of every line) → Design (character sheets, location plates) → Shots (beat list, coverage) → Frame-lock (compose each shot's opening frame as a still first) → Generate (one beat → review → one-variable delta if needed) → Chain (each next beat edits the previous approved frame, doesn't reinvent it) → Voice (to picture, never picture to voice) → Cut (hard cuts, overlapping audio across every edit, reactions) → Mix → Finish.
Continuity is not persuasion
A generated film can be perfectly continuous and still fail as marketing, sales, or creator work. Continuity tells you whether the world holds together. Persuasion asks a different question: what is the claim, where is the proof, and what belief or action is the viewer supposed to leave with?
Review order
Run the cheapest, most fundamental check first, and don't proceed past a stage that fails:
Anti-patterns to recognize and refuse
- Adding constraints until the performance has no room left — safety and craft compete for the same budget.
- Rewriting a whole prompt to fix one thing that's wrong.
- Naming a defect in a prompt in order to avoid it (this makes it more likely, not less).
- An emotion adjective with no named owner (in a two-person shot, it lands on both people).
- Directing the speaker and not the listener.
- Playing each generated clip end to end instead of cutting into it.
- Crossfades as a default connector between shots.
- Re-rolling a whole shot to fix one mispronounced word.
- Generating a second beat while the first is still unreviewed.
- A sales claim with no proof shot, only voiceover language trying to carry belief by itself.
- Creator-native pacing with copy that no actual person would say out loud.
- Claiming a clip is correct without having actually checked it.
- Passing a multi-person reference photo as a single-character identity reference.