How to Write Effective Prompts for AI Video Generation
A step-by-step guide to writing AI video prompts that work: subject and action, scene, camera, light, timing, sound and image-to-video, with a template and common mistakes.
A step-by-step guide to writing AI video prompts that work: subject and action, scene, camera, light, timing, sound and image-to-video, with a template and common mistakes.

A video prompt has to describe change over time. Say what moves, how it moves and what the camera does, then what the frame looks like.
Write it like a shot description for a film crew: one subject and one action, plus the setting, camera, light and style.
Keep each prompt to one shot. Several scene changes in one short clip usually come out muddled.
Use plain, physical verbs and concrete details. "She turns her head toward the window" beats "she looks contemplative."
Change one thing at a time between generations, so you can see what each word is doing.
An image prompt describes a single moment. A video prompt describes a few seconds of time, so the model has to work out what moves, in what order, and how the camera behaves while it happens. If you leave those things out, the model fills them in, and the usual result is a slow drift, a vague camera move or motion that doesn't match what you pictured.
Video models respond well to the terms film crews use to describe shots: shot sizes, camera moves, lens choices and lighting setups.
Lead with who or what the shot is about and what they're doing. Be specific about the subject, but put most of your words into the action, because that's what makes it a video.
Name one main subject: "an elderly fisherman," not "a man."
Give one clear action with a physical verb: "pulls a net over the side of the boat."
Add how it moves: speed, direction and weight, such as "slowly," "toward the camera" or "struggling with the weight."
Weak: "A fisherman on a boat."
Better: "An elderly fisherman in a yellow raincoat slowly hauls a heavy, dripping net over the side of a small wooden boat."
Next, say where and when it happens, in a few concrete details:
Location: "a small harbor in a gray North Sea town."
Time and weather: "early morning, light rain, mist on the water."
Background activity, if any: "gulls circling behind him."
Background motion is worth naming on purpose. If you don't want it, say the background is still. If you do, describe it, or the model may invent movement you didn't ask for.
Describe the camera in three parts.
Shot size tells the model how much of the subject to show: extreme wide, wide, medium, close-up or extreme close-up. Angle tells it where the camera sits: eye level, low angle looking up, high angle looking down, or overhead.
Use one move per shot, and name it with standard terms:
Static or locked-off: the camera doesn't move.
Pan or tilt: the camera turns left and right, or up and down, from a fixed spot.
Dolly or push-in: the camera travels toward or away from the subject.
Tracking: the camera moves with the subject, following behind, leading in front or traveling alongside.
Crane or drone: the camera rises, descends or flies over the scene.
Handheld: slight, natural shake, good for documentary or tense scenes.
Add speed if it matters: "a slow push-in" and "a fast push-in" produce very different shots.
Lens terms help with depth and mood. "Shallow depth of field" or "background softly out of focus" isolates the subject, "wide-angle lens" exaggerates space, and "rack focus from the net to his face" moves attention within the shot.
Say where the light comes from and what quality it has: "soft overcast daylight," "low golden sunlight from the left," "harsh overhead fluorescent light," or "lit only by a phone screen."
Then name the overall look in a few words: the color palette ("muted blues and grays"), the medium ("35mm film," "documentary footage," "stop-motion," "hand-drawn animation") and the era if it matters. Pick one style. Mixing several often produces a look that's none of them.
Most AI video clips run only a few seconds, which is enough for one shot. If you ask for "he hauls the net, then walks to the cabin, then the boat sails away," the model has to squeeze three shots into a few seconds and usually blends them into a confused one.
Write one prompt per shot instead, and cut the shots together afterward. That also gives you control over each shot's camera and timing.
Within a single shot, you can still have a short sequence. Use simple ordering words, and keep it to two or three beats:
"He pulls the net onto the deck, pauses to catch his breath, then looks up at the sky."
Words like "first," "then," "as" and "while" help the model place each beat. Long chains of actions tend to get dropped or merged, so if a beat matters, give it its own shot.
If your generator produces audio, describe it in the prompt like any other element. Name ambient sound ("rain on the deck, gulls, a distant engine"), specific sound effects tied to the action ("the wet slap of the net hitting the boards"), and music only if you want it. For dialogue, put the exact words in quotation marks and say who speaks and how: "He mutters, 'Not again,' in a tired voice."
If your generator doesn't make audio, leave sound out entirely so it doesn't take up the prompt.
When you start from an image, the image already sets the subject, setting, light and style. Repeating all of that in the prompt can fight the image or push the model to change it. Describe what should happen instead: the action, the camera move and any change in the scene.
Image-to-video prompt: "Slow push-in. The fisherman looks up from the net and squints into the rain as mist drifts across the water."
Make sure the motion you describe fits the image. If the image shows a closed door, asking for someone to open it and walk through works. Asking for a character who isn't in frame to appear tends to go badly.
Put the parts in this order and fill in what you need:
[Shot size and angle], [camera movement]. [Subject] [action, with speed and direction]. [Setting, time and weather]. [Light]. [Style and color]. [Sound, if supported].
For example:
Medium close-up at eye level, slow push-in. An elderly fisherman in a yellow raincoat hauls a heavy, dripping net over the side of a small wooden boat, then pauses to catch his breath. A gray harbor at dawn, light rain, mist on the water. Soft overcast daylight. Muted blues and grays, 35mm film look. Rain on the deck and the slap of the wet net.
Describing a still image: a beautiful frame with no action gets you a clip where almost nothing happens.
Abstract emotion words: "melancholy" or "epic" give the model little to animate. Describe what the viewer would see instead.
Too many subjects: each extra character is something else the model has to keep consistent. Start with one.
Conflicting directions: "static shot" and "sweeping drone view" in the same prompt cannot both happen.
Asking for readable text: lettering on signs and screens often warps as the camera moves. Add text afterward in an editor when it matters.
Writing what you don't want in the main prompt: "no people in the background" can put people in the background. Describe what you do want, or use a negative prompt field if your generator has one.
Treat each generation as a test. When a clip comes out wrong, work out which part failed (the action, the camera, the setting or the look) and change only that part of the prompt. If you rewrite everything at once, you won't know what fixed it.
A few fixes to try:
If the motion is too subtle, use stronger verbs and add speed: "quickly," "suddenly," "with force."
If the camera wanders, say "static camera" or name the one move you want.
If the subject changes appearance, make the description of the subject more specific, or start from a reference image.
If the clip looks generic, add one or two concrete, unusual details.
Keep a file of prompts that worked, with notes on what each part did.
Long enough to cover the subject, action, setting, camera and light, which is usually a short paragraph. Past that, extra detail gets ignored or crowds out the parts that matter.
Either the subject and action or the shot. You can lead with the subject and what it's doing, or open with the shot size and camera move. Both work, as long as the prompt names one subject, one action and one camera setup.
Usually because the prompt describes a scene rather than an action. Add a clear physical verb, a direction and a speed, and name a camera move.
Full sentences work better for most video models, because they show the order of events and what each detail belongs to. A list of keywords can't say what happens first.
Use the same detailed description of the character, word for word, in every prompt, and start each clip from the same reference image if your generator supports image-to-video.