Why the prompt is everything in AI video
Text-to-video (T2V) and image-to-video (I2V) models do not "understand" your idea โ they match your words to patterns in the film footage they were trained on. The prompt is the only instruction they receive, and small wording choices change what they generate. "A car on a road" and "a red vintage car driving along a coastal road at golden hour, tracking shot" produce two completely different clips from the same model.
This guide covers the structure that works: the six-part T2V prompt, the camera vocabulary video models respond to, how I2V prompts differ, and how to adapt one prompt for the model you are using.
T2V vs I2V: two modes, two different prompts
- T2V (text to video) โ you write the whole shot in words: subject, action, setting, camera, light, style. The model invents the appearance of everything.
- I2V (image to video) โ you supply a still image that fixes the look, and the prompt describes only what happens: the motion and any camera movement.
The difference matters because it changes what your words are for. In T2V the prompt is the entire storyboard. In I2V it is a motion note attached to an image you already chose.
The six-part T2V prompt formula
Every strong T2V prompt answers six questions, in roughly this order:
- Subject โ who or what is on screen, with concrete details (not "a person" but "a lighthouse keeper in a waxed coat").
- Action โ one main action per shot. Stacking three actions usually produces a blurred mixture of all of them.
- Environment & time โ where it is, what time of day, and the weather. Time of day drives the palette more than any style word.
- Camera โ shot size (close-up, wide...), angle (eye level, low angle...), movement (static, pan, dolly, handheld...). Naming even one of these stops the model from guessing.
- Lighting & color โ the look of the light: golden hour, blue hour, neon, candlelight; and a color direction like teal & orange or monochrome.
- Style & mood โ the visual language (cinematic, documentary, anime, noir) and the emotional tone (melancholic, tense, playful).
Worked example
Assembled from the six parts above:
"medium shot from the waist up, a lighthouse keeper in a waxed coat walking along a sea wall, waves breaking below, dusk with wind in the coat, tracking shot following the subject, blue hour light with a single warm lantern glow, cinematic film look, melancholic mood"
Count it: shot size, subject, action, setting, time, movement, light, style, mood โ nine elements, each clause doing one job, none of it decoration.
Camera vocabulary video models understand
Video models are trained on footage annotated with shot terminology, so the standardized film terms work. The words you will use most:
- Shot size โ extreme close-up, close-up, medium close-up, medium shot, medium wide shot, full shot, long shot, extreme wide shot, overhead shot, two-shot.
- Angle โ eye level, low angle, high angle, bird's-eye, worm's-eye, dutch angle, over-the-shoulder, point of view (POV).
- Movement โ static, pan, tilt, dolly in, dolly out, tracking, orbit, crane, handheld, zoom.
- Pacing โ slow motion, real time, fast motion, time lapse. Naming the pace is one of the most underused levers: "slow motion" alone can transform a clip.
You do not need to name all of them. One camera decision โ "tracking shot" or "handheld" โ often moves the result more than three style adjectives.
How I2V prompts work
In image-to-video mode the still image is already half the prompt. The image fixes the subject, the setting, the palette, and much of the style. Your text should only describe motion and camera:
- What happens: "the waves begin to roll in, the figure raises a lantern".
- How the camera behaves: "slow dolly in", "static, subject moves", "handheld follow".
- Tempo: "slow motion" or "real time".
Keep the description short โ one or two sentences is usually enough. If you re-describe the image's content in words, the model blends your text with the photo, and the result drifts from both. The AI Video Prompt Generator adds a "Following the input image" prefix in this mode and hides the categories the image already decides, such as shot size and composition.
Matching the prompt to the model
The major video models come from different labs: Kling (Kuaishou), Seedance (ByteDance), Hailuo (MiniMax), Wan 2.2 (Alibaba) and Vidu (Shengshu Technology), plus Sora 2 (OpenAI), Veo 3 (Google), Runway Gen-4 (Runway), Luma Dream Machine (Luma AI) and Pika 2.2 (Pika). Their prompting strategies differ in two practical ways:
- Language. The Chinese models are native in Chinese โ Chinese prompts read most naturally for them; the rest were trained mainly on English. Write the prompt in the language the model is strongest in.
- Quality phrase. Each model has a preferred finishing phrase โ Kling likes "cinematic quality, delicate light and shadow, rich detail", Pika 2.2 "cinematic, vibrant, smooth motion, crisp detail". Appending the right one makes the same selection read as a native prompt for each model.
The generator handles both: it matches the output language to the model (Chinese for the Chinese models on the Chinese site, English everywhere else) and appends each model's own quality suffix.
Seven mistakes that ruin AI video prompts
- Vague subject. "A person somewhere" leaves the model inventing the cast; be concrete about who or what is on screen.
- Multiple actions. Stacking three actions produces a blurred mix; one action per clip.
- Conflicting styles. "Documentary handheld" and "glamorous studio look" cancel each other; pick one visual language.
- No camera intent. Without shot size or movement the model guesses; declare it explicitly.
- Over-stuffing. A 200-word paragraph dilutes every instruction; 30-80 words of concrete descriptors usually beat a long essay.
- Mixed languages. Keep the prompt in one language; half-translated prompts confuse the model.
- Ignoring pacing. If you want slow motion or time lapse, say so; real time is the silent default.
Build prompts with structure
Instead of free text, you can pick options from a structured grid. The AI Video Prompt Generator covers 11 categories โ shot size, camera angle and movement, motion pacing, time and weather, lighting, color grade, style, mood, effects, and composition โ with reference images for every option, and assembles your single choice per category into one prompt tuned to the target model. It runs entirely in the browser: nothing you write is uploaded.