All guides

Craft 11 min August 2026

How to Brief an AI Video Project

What to ask, which storyline framework to pick, and the creative-brief template that keeps the first generation from being a wasted draft.

Japesh Singhal

Craft & production

Download skill

The most expensive mistake in an AI video project happens before the first generation — a brief that's thinner than it feels. This is the discovery pass worth running first: what to ask, which narrative framework to reach for, and the review habit that catches a weak prompt before it costs a generation credit instead of after.

01

Discovery, before anything else

Infer what you can from context; ask only what's genuinely missing. Group questions so it feels like being understood, not interrogated.

QuestionWhy it matters
OutcomeWhat should the viewer do or feel after watching? (Remember the brand, click, sign up, trust the product, laugh then feel relief — these need different films.)
AudienceWho is this for — one primary persona, not a demographic range?
PlacementWhere will this actually live? Aspect ratio and length follow from this.
Source materialWhat already exists — real footage, brand assets, a character idea only, or fully generated from scratch?
Who has to stay the same personAnyone recurring or visually important needs a locked identity reference before video generation starts, not stills improvised per scene.
Brand/productWhat must appear, and what must never appear?
TonePick 2–3 anchors, not a mood-board essay.
Hard constraintsWhat's absolutely off the table?
Reflect it back before drafting anything
Summarize in 3–5 sentences: "You're making a [format] for [audience] on [platform]. The viewer should [outcome]. Tone is [tone]. We'll [source approach] and avoid [constraints]." Then ask: does this match what you're trying to make? Catching a misread here costs a sentence; catching it after a draft costs the draft.
02

Pick a storyline framework — don't invent structure from scratch

FrameworkShapeBest for
3-act (Hook → Message → CTA)~5s hook, ~17s single idea, ~8s clear next step30-second TV or premium digital spots — one idea per spot, roughly 65–75 spoken words total if voiceover-led
5-shot performanceHook → Problem → Solution → Proof → CTAFast-paced social/performance ads — trim each generated clip to its 1–3 highest-signal seconds in the edit
Hook–Story–OfferAttention → compressed emotion or comedy → offerAnything from a 6-second bumper to a 30-second spot

One shot per generation call, assembled afterward — don't try to prompt an entire ad in a single generation. This is broad industry consensus, not an Ekly-specific quirk.

03

A creative-brief template

Copy, fill, and confirm with whoever's requesting the video before generating anything

Title — Objective (one KPI) — Audience (one persona: who, where, language) — Platform (ratio, length, sound-on vs. sound-off default) — Source (existing footage / full AI / hybrid) — Hook (exact first-2-seconds visual or line) — Key message (one sentence) — Proof/demo beat — CTA — Tone (2-3 words) — Visual style — Characters (names, roles, relationship) — Must show — Must not — Dialogue table (beat / speaker / exact line) — Deliverables (hero length, cutdowns, aspect ratios) — Constraints (broadcast/family-safe, legal/claims).

04

Review every prompt before you generate — a rubric, not a vibe check

Score each dimension Pass / Fix / Block — Block means don't generate until it's fixed.

CheckPass looks likeFix if...
Hook specifiedThe opening action/visual is first in the promptMove the hook to the top; name the first 2 seconds explicitly
One ideaA single gag, or a single line, or a single product momentSplit it into separate beats
Mute-readableThe story is clear with the sound off (for social)Add visual contrast or a physical gag instead of relying on audio
CTA separatedAny call-to-action lives in the edit, not the generation promptRemove URLs/on-screen text from the generation prompt itself
Brand kindnessThe people in the ad aren't mocked; the product reads as reliefReframe the copy
05

Helping someone who doesn't know their storyline yet

  • Anchor on the real situation, not a feature list. What actually started this — a real frustration, a piece of footage, an event?
  • Offer two spine options, don't leave it open-ended. A concrete choice between two named directions moves things forward faster than one more open question.
  • Steal a proven framework rather than inventing structure from nothing.
  • Write dialogue last. Lock the visual beats first — lines are easy to swap once the shots exist.
  • Promote important "background" people to real characters. If someone recurs across scenes, lock their identity reference before video generation, rather than letting each shot re-cast them.
"Make it premium" with nothing more specific? Propose a named menu, don't guess blind
A vague upgrade request answered with one guessed redo just restarts the same guessing loop when that guess gets rejected too. The fix that actually works: propose several distinctly named directions — each a one-line visual philosophy, not just "try again but nicer" — and let the person choose a lane, then push harder inside that one choice. This is the same principle used for a scene that isn't landing (see comedy timing) — a diagnostic ladder with named options, not one more guess.
06

The three-layer model for a recurring campaign

If a project has an identity (a character, a spokesperson) that has to hold across multiple videos, vary one layer at a time when iterating:

LayerWhat it coversVary it...
IdentityFace, outfit, voice — shouldn't change per variantRarely, and only deliberately
MessageHook, script, CTA, pacingPer variant — this is what A/B testing usually means
RenderingScene, camera, platform exportPer shot

If identity is drifting across shots, reduce how much the scene/camera is changing before rewriting the whole prompt — don't fix an identity problem by changing everything at once.