Just in

📋 AI Storyboarding: Script to Shot List to Screen (2026)

The five-stage AI storyboarding pipeline we run on every production: script to beats, boards, keyframes, clips — and exactly where each stage breaks.

Tomás Rivera

Tomás Rivera · Film Editor & AI Cinematography Writer

· 9 min read

✓ Fact-checked & production-testedBased on our own paid generations and published videos. Last reviewed 2026-08-12.How we test →
⚡ TL;DR — quick answers
Do I actually need a storyboard for AI video, or can I just prompt clip by clip?
You need one the moment a project has more than three shots that have to agree with each other — same character, same room, same time of day, a cut that has to match. Prompting clip by clip works for a single standalone shot. The second you need shot four to remember what shot two established, you're storyboarding whether you call it that or not; the only question is whether you do it on purpose, before you spend render credits, or discover the mismatch in the edit after you've already paid for the footage.
What's the difference between a storyboard and an animatic in an AI pipeline?
A storyboard is the still frames — one panel per shot, locking composition, character position and camera angle before anything moves. An animatic is those same panels cut together on a timeline with rough timing and a scratch voiceover or music bed, so you can watch the pacing of the whole sequence before you generate a single video clip. Skip the storyboard and you're guessing at composition. Skip the animatic and you're guessing at pacing — and pacing is the harder mistake to fix after the clips exist, because now you're re-rendering instead of re-arranging stills.
Should I build one big multi-panel storyboard image or generate each panel separately?
Build the grid as one image when character and setting need to hold across the sequence — a single generation sees all the panels at once and carries the same face, wardrobe and location through them far more reliably than nine separate calls to an image model that has no memory between requests. Generate panels separately once you've locked a reference (a headshot and a full-body shot) and just need one new composition — at that point the grid's consistency advantage stops mattering and single panels are faster to revise.
Cinematic AI video production illustration for: AI Storyboarding: Script to Shot List to Screen (2026)

Every AI video pipeline breaks at the same seam, and it isn't the model. It's the seam between "I have a script" and "I have a shot list" — the step most people skip because an image model will happily generate something gorgeous the moment you type a sentence at it, and gorgeous feels like progress. I spent twenty years on an editing bench before any of these models existed, cutting commercials and shorts from footage other people shot, and the one thing that never changed is this: you cannot fix in the edit what was never covered on the day. AI video has the exact same problem, just compressed into a single afternoon instead of a shoot week. Skip the storyboard, generate nine clips off nine improvised prompts, and you'll get nine beautiful shots that don't cut together — wrong axis, character's jacket changed color, the reaction shot doesn't match the line it's reacting to. This is the pipeline we actually run, stage by stage, borrowed almost wholesale from how documentary productions break a script before a single camera rolls.

Why documentary method, not demo-reel method

Most AI video advice comes from people showing off a single striking clip, which is a demo reel, not a production process. A demo reel doesn't have to hold together with the shot before or after it. A documentary does — and documentary production solved the "many shots, one coherent story" problem long before generative models existed, because a documentary crew can't afford to discover a missing shot after the interview subject has gone home. The method is blunt and it transfers almost unchanged: you don't shoot from a script, you shoot from a shot list, and the shot list comes from breaking the script into beats first. Every beat gets exactly the shots it needs — no more, because that's wasted budget, no fewer, because that's a hole you'll find in the edit. AI video removes the "the subject already left" constraint, but it doesn't remove the underlying discipline, and treating each clip as a free, independent roll of the dice is how a project quietly turns into thirty unrelated renders instead of one sequence.

The five-stage pipeline

Five numbered steps running script to beats, beats to shot list, shot list to a 3×3 grid, board to keyframes, and keyframes to clips.
The image model only enters at stage three; the two stages before it decide whether the sequence cuts together at all.

Stage 1 — Script to beats. Before any image gets generated, break the script (or the loose idea, if you're not scripting dialogue) into beats: the smallest units where something changes — a new piece of information, a new emotional state, a new location. A beat is not a shot. A beat is the reason a shot exists. Skipping straight from script to storyboard is where most people start prompting adjectives instead of blocking a scene, and it shows up later as a board that looks busy but has no throughline.

Stage 2 — Beats to shot list. Each beat gets translated into the actual shots that will cover it — typically a wide to establish, a medium or close to carry the emotional beat, and sometimes an insert or reaction to give the edit somewhere to cut. This is the step that determines whether your sequence has coverage. A single hero shot per beat is a slideshow with motion, not a scene, because a slideshow with motion can't survive a jump cut, a pacing fix, or a client asking for it ten seconds shorter.

Stage 3 — Shot list to storyboard grid. Now, and only now, do you touch an image model. This is where the site's own default process comes in: one storyboard sheet per sequence, generated as a single multi-panel image — we default to a 3×3, nine-panel grid — rather than nine separate one-off generations. A single generation call sees the whole sheet at once, so the same face, wardrobe and set carry across all nine panels the way they wouldn't if you fired off nine independent prompts to a model with no memory between requests. Every panel gets its camera angle, its framing, and — this is the part people skip — its exact place in the shot list, so panel four is understood as "the reaction to panel three," not a stand-alone pretty picture. GPT Image 2, which OpenAI shipped into the API on April 21, 2026, is the model we default to for this stage: it reasons about image structure before it renders, which matters more for a nine-panel narrative sheet than it does for a single hero image, because the model is effectively planning a small sequence, not just a picture.

Stage 4 — Storyboard to keyframes. Once a board is approved — and it has to be approved before this step, not adjusted mid-render — the panels that will actually become clips get pulled out, cleaned up, and locked with a character reference (one headshot, one full-body shot, the combination that survives a video model's interpretation better than a busy turnaround sheet — we go deep on exactly why in our character-consistency reference system). This is also where you decide, panel by panel, which shots are actually earning motion and which are static beats better served by a hold. Not every panel needs to become a moving clip; some carry more weight as a two-second still in the cut.

Stage 5 — Keyframes to clips, clips to cut. Each locked keyframe goes to a video model with a prompt that describes motion only — the board already locked the look, so re-describing appearance in the video prompt just invites drift. Model choice matters here more than most guides admit: we default to Seedance 2.5 when a character has to hold across many cuts because its reference stack is the deepest currently available, and we reach for Kling's character elements when it's one face and speed matters more than a deep reference set — our Kling 3.0 vs Seedance 2.5 breakdown has the fuller comparison, and how to actually get access to Seedance 2.5 if you're not on it yet. Then, before you spend another render credit, cut what you have into a rough animatic — even at wrong duration, even with placeholder audio — because pacing problems are cheap to catch in an animatic and expensive to catch after every clip exists.

The continuity rules that actually hold a sequence together

Table of three storyboard failures, axis violation, unmatched action and reference drift, showing what each looks like and the stage to catch it.
Reference drift is the expensive one, because spotting it after the clips exist means re-rendering rather than trimming.

Three things break more AI storyboards than any prompt-wording issue. First, axis violations — cross the 180-degree line between two consecutive panels and a character who was facing right is suddenly facing left with no cut logic to justify it; the fix is deciding your scene's axis before panel one, not after panel six looks wrong. Second, unmatched action — a hand mid-gesture in panel four has to land somewhere sensible by panel five, or the cut reads as a flinch instead of a motion. Third, and this is the one that costs the most money because you don't catch it until the clip stage: reference drift, where the same "character" prompt produces a subtly different face or wardrobe across panels because nothing was actually locked, just described. Here's the harshest line I'll write in this piece: a continuity failure you didn't catch on the board is not a small fix later — it's a full re-render, because you can't patch a mismatched face in the cut the way you can trim a frame. Catch it on the still, where it costs a regenerated panel, not a regenerated clip.

One cutting-room story, because it's the exact shape of the AI mistake and it earns its place here. Years ago I got a scene from a director who was so happy with one particular take of a walk-and-talk that he built the whole sequence around it — except he'd never shot the reverse for the other actor's coverage. In the edit there was nowhere to cut to when the second actor answered; every option either broke the axis or left a beat with no reaction shot. We ended up burying the fix in a slow push-in that hid the gap, and it worked, but only because film forgives a save like that once per scene. AI storyboards don't get the save. If your board doesn't have the reverse angle, there is no cutaway to hide behind — the model will happily generate a gorgeous shot for the beat you're missing, but it won't tell you that you needed it in the first place. That's the board's job, not the model's.

Storyboard software, or the DIY grid — which is worth it

Comparison table of Boords, Katalist and the DIY 3×3 grid, with columns for entry paid tier and what each adds to the workflow.
Paid tools buy shared review and version history; the DIY grid buys one fewer place for a character reference to drift.

Dedicated AI storyboard tools exist and some are genuinely good at the script-ingestion step. Boords added an AI storyboard generator with a free entry tier and paid plans starting around $39/month, and it handles script upload plus an automatic animatic pass — useful if your team needs a shared, editable board outside an image-model workflow. Katalist goes further, promising a full script-to-storyboard-to-video pipeline with character-consistency tooling and per-scene posing control, starting in the low tens of dollars a month for a solo plan. Both are worth trying if you want a dedicated interface with version history and client review built in.

What we actually run day to day is the DIY grid method above, because it stays inside the same model family we're already using for the keyframe and clip stages, and one fewer tool in the chain is one fewer place for a character reference to drift when it crosses a file format. If your team needs shared review links and client sign-off baked into the tool, pay for Boords or Katalist. If it's you and a render budget, the grid-first pipeline gets you to an approved board faster and cheaper — see our AI video cost breakdown for how storyboard-first budgeting compares to prompt-and-pray render costs, and our rundown of the mistakes beginners make covers most of the ways a skipped board shows up later as wasted credits.

By the numbers

Four number cards reading 9 panels per board, 2 character reference images, up to 50 Seedance 2.5 reference inputs, and about $39 a month for Boords.
The pipeline runs on a handful of fixed counts, so the cost of getting it right is discipline rather than software spend.
ItemStatusFigure
GPT Image 2 API launchConfirmed (OpenAI)April 21, 2026
Default storyboard gridSite standard3×3, nine panels per sequence
Character reference set for clip stageSite standardOne headshot + one full-body shot
Boords AI storyboard generator, entry paid tierConfirmed (Boords pricing page)~$39/month
Katalist storyboard-to-video, entry paid tierConfirmed (Katalist pricing page)Low tens of dollars/month
Seedance 2.5 reference inputsConfirmed (Higgsfield)Up to 50 multimodal references

How we verified: tool pricing traces to Boords' and Katalist's own pricing pages, loaded directly this week; the GPT Image 2 launch date traces to OpenAI's own announcement. The pipeline stages and continuity rules are our own production standard, run across every multi-shot sequence on this site.

The honest read: none of this is exotic. It's the same script-to-shot-list discipline documentary crews have used for decades, applied to a pipeline where the "camera" is an image model and the "shoot day" is a render queue. The teams getting clean, cuttable AI sequences aren't the ones with the best single prompt — they're the ones who wrote the shot list before they touched the model at all.

Frequently asked questions

Do I actually need a storyboard for AI video, or can I just prompt clip by clip?

You need one the moment a project has more than three shots that have to agree with each other — same character, same room, same time of day, a cut that has to match. Prompting clip by clip works for a single standalone shot. The second you need shot four to remember what shot two established, you're storyboarding whether you call it that or not; the only question is whether you do it on purpose, before you spend render credits, or discover the mismatch in the edit after you've already paid for the footage.

What's the difference between a storyboard and an animatic in an AI pipeline?

A storyboard is the still frames — one panel per shot, locking composition, character position and camera angle before anything moves. An animatic is those same panels cut together on a timeline with rough timing and a scratch voiceover or music bed, so you can watch the pacing of the whole sequence before you generate a single video clip. Skip the storyboard and you're guessing at composition. Skip the animatic and you're guessing at pacing — and pacing is the harder mistake to fix after the clips exist, because now you're re-rendering instead of re-arranging stills.

Should I build one big multi-panel storyboard image or generate each panel separately?

Build the grid as one image when character and setting need to hold across the sequence — a single generation sees all the panels at once and carries the same face, wardrobe and location through them far more reliably than nine separate calls to an image model that has no memory between requests. Generate panels separately once you've locked a reference (a headshot and a full-body shot) and just need one new composition — at that point the grid's consistency advantage stops mattering and single panels are faster to revise.

Which AI video model should turn my storyboard panels into clips?

Match the model to what the shot needs, not to whichever one is trending. Seedance 2.5 currently has the deepest reference-locking system of any consumer video model, which makes it the default when a character has to survive many cuts. Kling 3.0's custom character elements are faster to set up for a single recurring face. Neither model reads your storyboard panel as instructions on its own — you still have to write the shot's camera move, pacing and action into the prompt, the same way a DP reads a board and then decides how the camera actually moves to get there.

How many panels should one storyboard cover?

One panel per shot, not per beat — a single beat in your script often needs two or three shots (a wide, then a reaction, then an insert) to cut cleanly, and collapsing them into one panel is how you end up with a board that looks complete but has no actual coverage. Our default grid is nine panels, a 3×3 sheet, because it maps to roughly one clean sequence of a short-form piece without forcing you to split a single narrative beat across two separate boards.

The 5 best AI video finds, every week

New models, tested prompts, and what actually worked in our production — one short email a week. No spam, unsubscribe anytime.

Tomás Rivera

Written by Tomás Rivera

Film Editor & AI Cinematography Writer

Cut commercials and short films for two decades before AI video existed, and now grades every generator the way he graded dailies. Cares about continuity, coverage, and where the cut breaks — not demo reels.

#ai storyboarding#storyboard to video ai#script to storyboard ai#ai previz pipeline#shot list ai video

Keep learning