๐Ÿงฐ UtlKit

How to Write AI Video Prompts: T2V & I2V Structure, Camera Language & Examples

Why the prompt is everything in AI video

Text-to-video (T2V) and image-to-video (I2V) models do not "understand" your idea โ€” they match your words to patterns in the film footage they were trained on. The prompt is the only instruction they receive, and small wording choices change what they generate. "A car on a road" and "a red vintage car driving along a coastal road at golden hour, tracking shot" produce two completely different clips from the same model.

This guide covers the structure that works: the six-part T2V prompt, the camera vocabulary video models respond to, how I2V prompts differ, and how to adapt one prompt for the model you are using.

T2V vs I2V: two modes, two different prompts

  • T2V (text to video) โ€” you write the whole shot in words: subject, action, setting, camera, light, style. The model invents the appearance of everything.
  • I2V (image to video) โ€” you supply a still image that fixes the look, and the prompt describes only what happens: the motion and any camera movement.

The difference matters because it changes what your words are for. In T2V the prompt is the entire storyboard. In I2V it is a motion note attached to an image you already chose.

The six-part T2V prompt formula

Every strong T2V prompt answers six questions, in roughly this order:

  1. Subject โ€” who or what is on screen, with concrete details (not "a person" but "a lighthouse keeper in a waxed coat").
  2. Action โ€” one main action per shot. Stacking three actions usually produces a blurred mixture of all of them.
  3. Environment & time โ€” where it is, what time of day, and the weather. Time of day drives the palette more than any style word.
  4. Camera โ€” shot size (close-up, wide...), angle (eye level, low angle...), movement (static, pan, dolly, handheld...). Naming even one of these stops the model from guessing.
  5. Lighting & color โ€” the look of the light: golden hour, blue hour, neon, candlelight; and a color direction like teal & orange or monochrome.
  6. Style & mood โ€” the visual language (cinematic, documentary, anime, noir) and the emotional tone (melancholic, tense, playful).

Worked example

Assembled from the six parts above:

"medium shot from the waist up, a lighthouse keeper in a waxed coat walking along a sea wall, waves breaking below, dusk with wind in the coat, tracking shot following the subject, blue hour light with a single warm lantern glow, cinematic film look, melancholic mood"

Count it: shot size, subject, action, setting, time, movement, light, style, mood โ€” nine elements, each clause doing one job, none of it decoration.

Camera vocabulary video models understand

Video models are trained on footage annotated with shot terminology, so the standardized film terms work. The words you will use most:

  • Shot size โ€” extreme close-up, close-up, medium close-up, medium shot, medium wide shot, full shot, long shot, extreme wide shot, overhead shot, two-shot.
  • Angle โ€” eye level, low angle, high angle, bird's-eye, worm's-eye, dutch angle, over-the-shoulder, point of view (POV).
  • Movement โ€” static, pan, tilt, dolly in, dolly out, tracking, orbit, crane, handheld, zoom.
  • Pacing โ€” slow motion, real time, fast motion, time lapse. Naming the pace is one of the most underused levers: "slow motion" alone can transform a clip.

You do not need to name all of them. One camera decision โ€” "tracking shot" or "handheld" โ€” often moves the result more than three style adjectives.

How I2V prompts work

In image-to-video mode the still image is already half the prompt. The image fixes the subject, the setting, the palette, and much of the style. Your text should only describe motion and camera:

  • What happens: "the waves begin to roll in, the figure raises a lantern".
  • How the camera behaves: "slow dolly in", "static, subject moves", "handheld follow".
  • Tempo: "slow motion" or "real time".

Keep the description short โ€” one or two sentences is usually enough. If you re-describe the image's content in words, the model blends your text with the photo, and the result drifts from both. The AI Video Prompt Generator adds a "Following the input image" prefix in this mode and hides the categories the image already decides, such as shot size and composition.

Matching the prompt to the model

The major video models come from different labs: Kling (Kuaishou), Seedance (ByteDance), Hailuo (MiniMax), Wan 2.2 (Alibaba) and Vidu (Shengshu Technology), plus Sora 2 (OpenAI), Veo 3 (Google), Runway Gen-4 (Runway), Luma Dream Machine (Luma AI) and Pika 2.2 (Pika). Their prompting strategies differ in two practical ways:

  • Language. The Chinese models are native in Chinese โ€” Chinese prompts read most naturally for them; the rest were trained mainly on English. Write the prompt in the language the model is strongest in.
  • Quality phrase. Each model has a preferred finishing phrase โ€” Kling likes "cinematic quality, delicate light and shadow, rich detail", Pika 2.2 "cinematic, vibrant, smooth motion, crisp detail". Appending the right one makes the same selection read as a native prompt for each model.

The generator handles both: it matches the output language to the model (Chinese for the Chinese models on the Chinese site, English everywhere else) and appends each model's own quality suffix.

Seven mistakes that ruin AI video prompts

  • Vague subject. "A person somewhere" leaves the model inventing the cast; be concrete about who or what is on screen.
  • Multiple actions. Stacking three actions produces a blurred mix; one action per clip.
  • Conflicting styles. "Documentary handheld" and "glamorous studio look" cancel each other; pick one visual language.
  • No camera intent. Without shot size or movement the model guesses; declare it explicitly.
  • Over-stuffing. A 200-word paragraph dilutes every instruction; 30-80 words of concrete descriptors usually beat a long essay.
  • Mixed languages. Keep the prompt in one language; half-translated prompts confuse the model.
  • Ignoring pacing. If you want slow motion or time lapse, say so; real time is the silent default.

Build prompts with structure

Instead of free text, you can pick options from a structured grid. The AI Video Prompt Generator covers 11 categories โ€” shot size, camera angle and movement, motion pacing, time and weather, lighting, color grade, style, mood, effects, and composition โ€” with reference images for every option, and assembles your single choice per category into one prompt tuned to the target model. It runs entirely in the browser: nothing you write is uploaded.

Related Tools

Frequently Asked Questions

What is a text-to-video (T2V) prompt?

A T2V prompt is the complete text description a text-to-video model turns into a clip. Because the model sees no image, the prompt must specify the subject, the action, the environment, the camera (shot size, angle, movement), the lighting, and the style โ€” everything a storyboard would normally contain.

How is an image-to-video (I2V) prompt different?

In I2V the input image already fixes the appearance, the setting, and the palette. The prompt only describes what happens: the motion, the pacing, and any camera movement. Re-describing the image's content in words usually makes the model drift away from the picture.

What should an AI video prompt include?

Six parts: subject, one main action, environment with time of day and weather, camera (shot size, angle, movement), lighting and color, and style with mood. Keep it to roughly 30-80 words of concrete descriptors, in a single language.

How long should an AI video prompt be?

Longer is not better. Around 30-80 words of concrete, unambiguous descriptors is the sweet spot: enough to fix the shot, short enough that no detail gets diluted. A 200-word paragraph usually performs worse than a tight one.

Do I need a different prompt for each model (Kling, Sora 2, Veo 3)?

The core description can stay the same, but two things should be adapted: the language (Chinese models like Kling, Seedance, Hailuo, Wan and Vidu are trained to respond to Chinese prompts; the rest read English best) and the finishing quality phrase, which differs per model. A structured generator handles both automatically.

Related Articles