🌀 40 Image-to-Video Motion Prompts That Won't Wreck Your Still
40 motion-only image-to-video prompts for Seedance, Kling and Veo — ambient motion, subject actions, camera drift, atmosphere — and where each one breaks.
Amara Osei · Prompt Engineer & Workflow Writer
· 8 min read
⚡ TL;DR — quick answers
- What makes a prompt 'motion-only' for image-to-video?
- It describes nothing the still already shows — no subject, no wardrobe, no setting. It states exactly one thing: what moves, how it moves, and for how long. Every prompt in this library follows that rule on purpose. The moment a prompt redescribes the person or the room, the model starts negotiating between your text and your pixels, and that negotiation is where identity drift comes from.
- Why do image-to-video prompts fail even when the still looks perfect?
- Two variables, tested separately: motion count and motion size. Stack more than two moves in one prompt and the model starts averaging them into a blur. Ask for a move larger than the still can physically support — a full body turn from a headshot, a room reveal from a close-up — and it either warps the geometry or ignores half the instruction. Both failures trace back to the source image having no data for what you asked it to invent.
- Do the same motion prompts work across Seedance, Kling and Veo?
- The grammar transfers — one move, one clear verb, a duration. The tolerances don't. Seedance 2.5 holds ambient motion the longest before drifting. Kling's motion brush beats plain text when the move needs to stay inside a hand-drawn region. Veo 3.1 is strict about mixing a starting image with reference images in the same call — pick one or the other. Same 40 prompts, three different failure ceilings.

Forty prompts, one variable isolated per category, run against a fixed still per test: same source image, three renders per prompt, across Seedance 2.5, Kling 3.0 and Veo 3.1. The rule that survived all three models: the image owns identity, the prompt owns motion — nothing else. The moment a prompt redescribes what the still already shows, that's where drift starts.
Here's the source tutorial the batch was built against, for the raw Seedance workflow:
The one rule
Image-to-video is not text-to-video with a reference attached. The still already answers "who" and "where." Your prompt answers exactly one question: "what moves, how, and for how long." Every prompt below is built to that spec — subject or camera action, one modifier, an optional duration. No wardrobe notes, no lighting redescription, no scene-setting. If your current prompts open with a sentence describing the person or the room, that's the first thing to cut before you touch anything in this library — see the beginner mistakes we log most often for the rest of that pattern.
One vendor-specific gotcha worth knowing before you start: Veo 3.1 will not accept a starting image and reference images in the same generation call — you commit to one mode or the other before you ever get to the prompt. Seedance vs. Veo covers where else the two diverge on image inputs.
Test conditions
Same still per category, three runs per prompt, Kling 3.0 vs. Seedance 2.5 as the two-model baseline, Veo 3.1 spot-checked on the ones that survived both. A run counts as failed if the subject's face, hair color or clothing shifted from the source frame, or if the requested motion didn't read as the dominant action in the clip. Fail counts below are our own batch, not a vendor benchmark — treat them as a starting prior, not a guarantee on your still.
01–10: Ambient motion
The lowest-risk category and the one to reach for when you don't care about the exact final frame — hair, water, flame, dust. One micro-motion, nothing else moves. Fail rate across the batch: roughly 1 in 10, almost always from a still where the "still" element (candle, curtain) was already slightly blurred and the model had no clean edge to animate against.
v01 — Hair drift
Loose strands of hair lift and settle once in a light breeze. Nothing else moves.
v02 — Surface ripple
Small ripples spread across the liquid's surface, starting from the center. Cup stays still.
v03 — Steam
Steam rises off the cup in slow, curling wisps for 3 seconds. Cup does not move.
v04 — Background leaves
Leaves on the tree behind the subject tremble in a light wind. Subject holds still.
v05 — Curtain sway
The curtain behind the window sways inward once, then settles back.
v06 — Candle flicker
The candle flame flickers and leans left, then rights itself over 2 seconds.
v07 — Dust motes
Dust motes drift slowly through the shaft of window light. No other motion.
v08 — Natural blink
Subject blinks twice at a relaxed pace. No head or body movement.
v09 — Fabric hem
The hem of the dress sways once as a breeze passes through frame.
v10 — Tank bubbles
Bubbles rise steadily from the aerator in the tank behind the subject.
11–20: Subject actions
One deliberate, physically small action per prompt — the hand does one thing, the head does one thing. This is the category where redescribing the subject does the most damage, because the model already has a face and a body to match; adding text about either just gives it permission to "fix" something. Fail rate: roughly 1 in 4, almost all identity drift on actions that moved the face out of three-quarter view — see our consistent-character notes for the wider pattern.
v11 — Head turn
Subject turns their head from camera-left to center over 2 seconds. Expression unchanged.
v12 — Reach and lift
Hand reaches into frame, picks up the mug by the handle, lifts it one inch.
v13 — Slow smile
A smile forms gradually over 3 seconds. Eyes crease slightly at the corners.
v14 — Rise from chair
Subject rises from the chair, weight shifts forward first, then up.
v15 — Page turn
Hand turns one page of the open book. Page settles flat.
v16 — Sip and set down
Subject lifts the cup, takes one sip, sets it back in the same spot.
v17 — Adjust glasses
Subject pushes their glasses up the bridge of their nose with one finger.
v18 — Glance back
Subject glances back over their left shoulder, then returns to facing camera.
v19 — Push door open
Hand pushes the door open from the near side. Door swings to roughly 45 degrees.
v20 — Step into frame
Subject takes two steps forward from the left edge and stops centered.
21–30: Camera drift only
Subject stays frozen; the camera does the work. This is the cleanest category for testing how many moves a model can hold, because there's only one moving part to grade. One move: near-perfect across all three models. Two sequential moves, timed: still solid. Three moves in one prompt: every model in the batch started blending them — a push-in that also drifted left also tilted, with none of the three reading clearly. Keep it to one, two at most.
v21 — Slow push-in
Camera pushes in slowly to a medium close-up on the subject's face. Subject does not move.
v22 — Lateral dolly
Camera drifts left to right at a steady pace. Subject stays static in frame.
v23 — Pull-back reveal
Camera pulls back slowly from close-up to reveal the full room.
v24 — Tilt up
Camera tilts up from the subject's hands to their face over 2 seconds.
v25 — Partial orbit
Camera arcs a quarter turn around the subject, holding the same distance.
v26 — Rack focus
Focus racks from the foreground object to the subject's face. Camera stays static.
v27 — Low rise
Camera drifts a few inches upward from a low angle, like a slow rise off the floor.
v28 — Parallax pass
Camera moves right past a foreground object, revealing the subject behind it.
v29 — Handheld settle
Camera has a subtle handheld sway for 1 second, then settles to steady.
v30 — Crane down
Camera descends slowly from a high angle down to eye level.
31–40: Atmosphere and environment
Motion lives in the background — weather, light, particles — and the subject is untouched by design. This is the highest ceiling for "cinematic" payoff per word of prompt, and the one place stacking two atmospheric effects (say, fog plus a light shift) actually held up in testing better than stacking two camera moves did. Fail rate: roughly 1 in 6, mostly foreground/background bleed where "background" fog crept across the subject's face.
v31 — Fog roll
Fog rolls slowly across the ground in the background. Subject unaffected.
v32 — Light shift
Light shifts from cool blue to warm gold over the duration, as if a cloud passes.
v33 — Rain starts
Rain begins falling in the background, drops visible against the dark building facade.
v34 — Snow drift
Snowflakes drift down slowly in the background. None touch the subject.
v35 — Skyline lights
Background city lights flicker on one by one across the skyline.
v36 — Wind through grass
Tall grass in the background bends and ripples in a steady wind.
v37 — Fireflies
Fireflies blink on and off in the dark tree line behind the subject.
v38 — Cloud shadow
Clouds drift across the sky in the background, casting a moving shadow on the ground.
v39 — Neon flicker
The neon sign in the background flickers once, then holds steady.
v40 — Distant haze
A faint dust haze drifts across the horizon in the far background.
What actually breaks the still
Four failure modes showed up across all 40 prompts and all three models, in order of frequency:
- Redescription drift. Any clause that restates the subject — hair color, clothing, expression — gives the model room to "correct" it. Cut it. Start every prompt at the verb.
- Scale mismatch. Asking for motion the still has no data for — a hand reach when the hand isn't in frame, a room reveal from a tight close-up — produces warped geometry more often than it produces the requested shot. If the pixels aren't there, generate a new still instead of asking the model to invent them.
- Stacked moves. Two sequential, timed moves survive. Three simultaneous moves blend into unreadable motion on every model tested, not just one — this isn't a Seedance-specific or Kling-specific limit.
- Background bleed. Atmospheric effects described as "background" without a boundary clause occasionally crossed into the subject. Adding "subject unaffected" or "subject stays static" as a closing clause cut this failure close to zero in the retest.
None of these are model bugs so much as prompt debt — text asking for something the image can't support. Fix the prompt before you burn a second render.
Building this into a workflow
If you're choosing which model to run this library against, Kling 3.0 vs. Seedance 2.5 and Kling vs. Seedance both break down where each one's motion tolerance actually sits, and Seedance 2.5 vs. Veo 3.1 is the one to read before you commit to Veo's single-image-or-references choice. If you're still assembling your source stills, how to access Seedance 2.5 covers the current entry points, and our consistent-character notes are the companion piece for keeping the same face across a whole sequence of these prompts rather than one isolated clip.
Run one variable at a time, log which prompt number failed and why, and keep the failure — it tells you more about your source still than the pass does.
Frequently asked questions
▸What makes a prompt 'motion-only' for image-to-video?
It describes nothing the still already shows — no subject, no wardrobe, no setting. It states exactly one thing: what moves, how it moves, and for how long. Every prompt in this library follows that rule on purpose. The moment a prompt redescribes the person or the room, the model starts negotiating between your text and your pixels, and that negotiation is where identity drift comes from.
▸Why do image-to-video prompts fail even when the still looks perfect?
Two variables, tested separately: motion count and motion size. Stack more than two moves in one prompt and the model starts averaging them into a blur. Ask for a move larger than the still can physically support — a full body turn from a headshot, a room reveal from a close-up — and it either warps the geometry or ignores half the instruction. Both failures trace back to the source image having no data for what you asked it to invent.
▸Do the same motion prompts work across Seedance, Kling and Veo?
The grammar transfers — one move, one clear verb, a duration. The tolerances don't. Seedance 2.5 holds ambient motion the longest before drifting. Kling's motion brush beats plain text when the move needs to stay inside a hand-drawn region. Veo 3.1 is strict about mixing a starting image with reference images in the same call — pick one or the other. Same 40 prompts, three different failure ceilings.
▸How many camera moves can I stack in one image-to-video prompt?
One, cleanly. Two, if they're sequential and timed ('push in for 2 seconds, then settle'). Three is where every model in this test started blending moves into unreadable motion, regardless of vendor. If a shot needs three distinct moves, that's three prompts and an edit-room cut, not one generation.
▸What's the single most common image-to-video mistake?
Re-describing the still. Writers open with 'a woman in a red coat stands by a window' when the image already is that. The model then has to decide whether your text is a correction or a restatement, and it sometimes 'corrects' features that were never wrong — hair color shifts, coat changes shade. Drop the redescription. Start the prompt at the verb.
The 5 best AI video finds, every week
New models, tested prompts, and what actually worked in our production — one short email a week. No spam, unsubscribe anytime.

Written by Amara Osei
Prompt Engineer & Workflow Writer
Treats prompting as an experiment, not an art: isolate one variable, run the batch, keep the receipts. Her prompt libraries ship with failure rates, not just the wins.
Explore these topics
Every guide, comparison and prompt library we have on each.





