The single most common way an AI-generated film falls apart isn't a bad shot — it's a shot that's individually fine but doesn't match the one before it. A location's layout shifts. A logo redraws itself slightly differently each time. A prop's exact position drifts across a beat chain one small loss at a time until, five shots later, the world clearly isn't the same world anymore. None of this is inevitable — it's a discipline problem with a specific, learnable fix.
Lock your world with a master plate, then re-seed every shot from it
The technique
Before generating any shot in a recurring location, generate one master reference "plate" — a still image that locks the room, the set, or the object arrangement exactly as it should look. Then, for every later shot that revisits that location: re-seed the generation from the plate itself, not from whatever the previous shot happened to produce. This is hard rule 3 in the the 14 hard rules — lock the world as pixels; prose cannot hold a place across shots.
A real, observed example: a location plate correctly locked four distinct objects in a specific floor arrangement. A later shot, generated by extending an earlier shot's own frame rather than referencing the plate directly, quietly dropped to showing only one object — the count and arrangement the plate specified were both lost, and nobody caught it until someone looked closely.
The fix, as a checklist
- Every shot that shows a previously-locked location passes the original plate as an explicit reference — not whichever shot happened to come immediately before it.
- A previous shot's own output is a fine seed for the next shot's specific composition (the motion, the exact framing) — it's just never a substitute for the plate as the source of truth for what the location actually contains.
- Before calling a shot "done," diff it against the plate's locked facts specifically — object count, arrangement, position of background details — not just against the shot right before it.
When a background detail's exact partial state keeps failing
If a background object's precise partial visibility (distant, small, partly hidden) keeps coming back wrong no matter how many times you regenerate, stop fighting for exact partial precision and fully occlude it instead — put something else between the camera and the object so it isn't partially visible at all. A background element that's completely hidden reads as a deliberate framing choice; one that's almost-but-not-quite right in the background reads as an error every time.
Adding or removing an object without breaking the shot
Don't prompt a removal directly
If you need the same frame minus one object — a before-and-after, an item taken away, a repair — don't ask the model to "remove" it. Naming the thing you want gone ("no pin at the shoulder") is still a positive cue for that exact object, and a generative model will often re-include or half-include what you just told it to exclude.
Treat it as an edit, not a generation. Take the approved frame and edit it directly — inpaint, retouch — so everything except the one change is pixel-identical by construction. Then animate the edited still if it needs to move.
The re-parenting trick
Before inpainting a removal, check whether an earlier version in the same production chain already lacks the object. If it does, branch from that earlier version and add back everything that changed since (minus the one thing you need gone), instead of trying to remove the object from the current version. Additions are consistently more reliable than removals — so turning a removal into "add everything except X, starting from a version that never had X" tends to succeed where a direct removal attempt fails.
A real, observed example: an identity reference sheet needed a logo patch taken off one sleeve. Asking the model to remove it from the current version failed, the way direct removals tend to. But an earlier version of the same sheet — created before that patch had been added in the first place — already lacked it. Re-adding everything else that had changed since (a separate pocket badge) onto that earlier version worked on the first attempt.
The three kinds of reference image
Not every reference image is doing the same job. Mixing up the roles is a common, invisible cause of a shot that "should" have worked but didn't.
Brand marks need their own reference — every time
A location plate can correctly lock a room's geometry and still leave the model inventing its own version of a wall logo — wrong shape, wrong colors, a different wordmark entirely — if the brand asset isn't separately included as its own reference in that same generation call. The brand asset has to be in the reference set every time the mark is visible; it doesn't carry over automatically from an earlier shot that happened to get it right.
Flatten a photographed logo before using it as a reference
If the only asset you have is a phone photo of a physical printed sign or logo — angled, glossy, with a visible shadow — feeding that photo directly as a reference tends to produce a photo of the logo, not the logo's actual artwork: the right colors, but the internal shape subtly reorganized into something that only resembles the mark. Flatten the photo first — deskew it, remove the background and any shadow — and reference that clean, flat asset instead. Treat this as a standing pre-production step, done once before the first shot that needs the mark, the same way you'd generate a location plate ahead of time.
When more retries won't fix it
Two specific, recognizable situations where continuing to reprompt makes things worse, not better — and what to do instead.
A gesture that keeps escalating instead of converging
Precise, mutual two-person gestures — a synchronized nod, a shared glance, a moment of quiet acknowledgement between two characters — are a genuinely harder target than a single character's action, and reworded attempts on this category tend to get worse, not closer.
A real, observed sequence, on one shot: a request for a small head gesture between two people produced a hand-to-chin "thinking" pose. Adding negations ("no arm raise") produced an actual bow — because naming the gesture to ban it functioned as a cue for it. The next reword produced open mouths that read as mid-speech. Switching to a different model entirely produced two people walking toward each other "like a romance scene." Four attempts, four different wrong readings of the same handful of words.
What actually worked: stop animating the gesture at all. Hold a still frame of the approved two-shot for the beat's duration instead of generating motion for it, then cut. A brief, unforced moment between two people reads fine as a photograph — there's no "what happens between these two people" for the model to invent wrong, because nothing has to happen.
When the still-hold itself doesn't fully solve it
Even a still hold can feel like a dead pause if the beat's whole job was supposed to be a small moment of connection. When a beat's content keeps failing across both reworded prompts and a switch to a different model, that's real information: the specific thing being asked for may be past what current models can depict reliably — not that the right words haven't been found yet. At that point, the fix is changing what the beat actually shows — same story position, same idea, different concrete content — not a sixth attempt at the same premise. In the case above, the real fix wasn't a better still or a fifth model; it was rewriting that silent moment into a short dialogue exchange instead.