"Body movement doesn't feel real." "Face reactions are poor." "Too posed, too model-like." These complaints share one root cause: the prompt described a state instead of an action, and a model given a state renders the appearance of that state — which is precisely what bad acting looks like. The fix isn't more adjectives. It's verbs, named targets, and the same discipline a director gives a real actor on set.
Shot grammar, transitions, and rhythm in AI-generated film
One film, one idea
Before generating anything, write one sentence that is not the product and not the premise. A premise is the situation; an idea is what the film is about. If you can only state the premise, you have a sketch — a string of gags that could run in any order — not a film.
Find the idea already sitting in the material. The strongest directorial move is often latent and unused. A product that promises "your footage doesn't have to be a mess" can simply never show the footage — cut to a camera viewfinder, frame the chaos as footage being shot, reveal the same moment as a beautifully finished cut at the end. That's a directorial idea that costs nothing extra to generate — it's a choice about framing, not a new asset.
Camera as point of view, not container
Assign one motivated camera move per beat:
Uniform focal length across every shot is one of the loudest tells that nobody directed the film. Vary lens choice deliberately — wide for the chaotic ensemble beat, normal for the exchange, long reserved for the one moment that should feel like a change of key.
Transitions are grammar, not decoration
A dissolve means time passed or this is a memory — it does not mean "here's the next shot." Comedy in particular lives on the abruptness of hard cuts; the cut is the timing.
Rhythm: the difference between assembled and edited
Set your target average shot length from the genre, not a fixed rule — comedy wants short cuts, drama wants longer holds. *The failure to avoid isn't slowness — it's unchosen length.* If generated clips are all a similar length and you play them end to end, you land at an average shot length by accident, and that reads as slow, flat, and machine-made even when the shots themselves are fine. Ask of every hold: did you choose this duration, or did the generator?
One generation is not one shot. Every generated clip should yield two or three cuts — the setup, the reaction, the button — intercut with other coverage. This costs nothing extra and multiplies your effective shot count.
Use J-cuts and L-cuts on every cut. Every generated clip is made in isolation, so cutting picture and sound together at the exact same frame advertises the seam. Overlapping audio across the cut welds independently generated shots into one continuous space — typically a quarter to half a second of offset. This is the highest-impact, lowest-cost move available in the edit, and it also hides weak lip-sync, because you're off the mouth when it matters.
Shot grammar across the film
- Vary shot size deliberately — wide to establish, mediums for the exchange, one close-up reserved for the punchline.
- Plan coverage, not just clips — for every dialogue beat: the speaker, the listener, and one detail insert.
- Hold the 180-degree line. Generative models don't preserve screen direction between independent generations — you have to name the camera side in every single prompt ("camera on their left, over the shoulder"), or both characters end up looking the same direction and the scene reads subtly wrong to everyone and legible to no one.
- A still is a shot type, and the cheapest one you have. A freeze reveals "you're looking at material, not a moment" — the strongest and cheapest way to signal that. A hold on a still buys a beat for sound to do something. Stills carry no lip-sync risk and no invented speech.
Comedy timing in AI-generated video
Comedy is structure, not jokes
A comic beat has four movements, same shape as the whole film: setup (establish normal, who wants what), escalation (the same want, three times, each worse), turn (the thing everyone forgot walks in), button (one image or line that lands the idea).
- Rule of three, with a break. Two examples establish a pattern; the third has to violate it. Three parallel complaints is a list. Two complaints and a reversal is a joke.
- Escalate the stake, not the volume. Louder isn't funnier — the second beat should cost the character more than the first.
- The turn must be inevitable and unexpected. Plant it in the setup so the audience feels clever for not seeing it, not cheated.
- The button is an image first. If the ending only works because someone says the line, you have an announcement, not an ending.
Where the joke actually lives: the reaction, not the line
A line delivered to a frozen, non-reacting listener dies regardless of how well it's written or performed. This is also the cheapest coverage available — reaction shots need no dialogue, no lip-sync pass, and often no generated audio at all.
- Comedy is status. Every funny scene is someone's status moving. Direct every performance as a status move, not as an emotion.
- Specificity beats generality, always. A specific detail is the joke; a generic one is a complaint. When a line feels flat, it's almost always too general — don't make it louder, make it more specific.
- Play the opposite. The character who's furious plays it wounded and dignified. The character who's smug plays it generous. Comedy comes from the gap between what someone feels and what they perform.
- The straight face is the engine. Nobody in the film should know it's a comedy. The moment a performance signals "this is the funny bit," the joke is gone.
When the reviewer says it's not funny
Diagnose in this order — the fix is almost never a new line:
The scene isn't landing
Symptom — A reviewer says "it's just not funny"
- Is there a reaction shot? Add one before touching the writing. Is the pace too slow? Cut a second from every beat and re-watch. Is a detail too general? Make one noun specific. Is the performance signaling? Play it straighter, not bigger. Is the structure just a list? Break the pattern on the third beat. Is the camera doing anything? Add one push-in on the punchline. Only after all of the above: rewrite the line.
Emotional and dramatic structure in AI-generated video
Where the feeling lives
Comedy compresses; emotion delays. A joke wants its punch word early enough to land and late enough to surprise. A feeling wants the beat after the event — the pause where someone absorbs it. Cut a comic beat one frame early; let an emotional beat run one beat past where you'd normally cut.
The audience does the work, or there's no feeling. Emotion comes from the gap between what's shown and what's understood — show the hand, not the tears; the empty chair, not the grief. Stating the feeling relieves the audience of producing it, and a stated feeling is information, not experience.
Reveal architecture: plant, carry, pay
The test of a good plant: it's useful the first time, on its own terms. A detail whose only job is to be paid off later reads as a setup and kills its own surprise.
Decide, per scene, whether the audience knows more than the character, less, or the same — and decide it deliberately. Knowing more produces dread and tenderness (we watch someone walk toward something they can't see). Knowing less produces mystery. Knowing the same produces identification, but has to borrow its tension from what the character wants.
Directing someone who never speaks
Silence isn't the absence of performance — it's a performance with a different instrument. Give them an unbroken want (so stillness reads as containment, not vacancy), one thing to look at, and something to not do — a hand that stops short, a step not taken. A withheld action is legible and specific. And let another character speak into their silence — silence is only loud beside sound.
When the reviewer says it doesn't move them
Same discipline as comedy's ladder, different rungs — the fix is almost never "add more emotion," because adding emotion is usually what produced the flatness in the first place.
The scene isn't landing emotionally
Symptom — A reviewer says "it just doesn't move me"
- Is anything withheld? If the film states its feeling outright, remove the statement and see what survives — usually the scene gets better immediately. Does the audience know something the character doesn't? If everyone knows the same things, there's no tension to spend. Is there a beat after the beat — the moment someone absorbs what just happened? If it isn't in the cut, it isn't in the film. Is it too fast? Emotional beats cut at comic pace read as glib. Is anyone alone in frame? Feeling needs a single shot — two-shots distribute attention, and distributed attention isn't intimacy. Does the music arrive before the feeling has earned it? Delay it, or drop it entirely at the turn. Only after all of these: change the actual story, not just the performance.
Directing body language, face, and delivery
Why generated performance reads dead
The root error is almost always the same: the prompt describes a state instead of an action.
Verbs, not adjectives
Scope every direction to one named person. A performance note with no owner gets applied to everyone in frame — write "Character A: [note]" and give everyone else in the shot an explicit instruction too, even if it's just "does nothing."
Body language: status is physical
In any scene where two people want different things, performance is status, and status is posture.
Give every character one physical signature and keep it across the film — a hand that always returns to the same place, a habit of straightening something. Signatures survive generation better than descriptions of personality. Weight is the tell — "standing" produces a mannequin; "weight on the back foot, one shoulder lower" produces a person.
The face
- Micro before macro. The real expression is the flicker before the composed one. "For a moment his face doesn't move at all, then he smiles" is a performance; "he smiles politely while annoyed" is a label.
- Delay is everything. Reactions that land on the beat are fake — real ones land a fraction late.
- The eyes lead, and always need a named target. Unspecified eyes go dead or wander.
- Blink, swallow, breathe. One involuntary action per shot separates a person from a rendering.
Listening is a performance
The most common failure in generated dialogue scenes: only the speaker was directed. The listener stands frozen while a line is delivered at them. In comedy, the reaction is the joke; in drama, the listener carries the meaning. Direct the listener in the same prompt as the speaker — and then shoot the listener's reaction separately, since it needs no dialogue and no lip-sync pass.
Register by genre
Performance scale is genre, and the mismatch between them is why films feel tonally wrong even when everything is technically correct: