Why the prompt is everything in AI images
Text-to-image (T2I) models do not "understand" your idea — they match your words to patterns in the billions of images they were trained on. The prompt is the only instruction they receive, and small wording choices change what they paint. "A dog in a park" and "a scruffy terrier mid-leap after a tennis ball, dappled sunlight through the trees, 85mm lens, shallow depth of field" produce two completely different pictures from the same model.
This guide covers the structure that works: the eight-part T2I prompt, the visual vocabulary image models respond to, how to adapt one prompt for the model you are using, and the mistakes that flatten results.
The eight-part T2I prompt formula
Every strong T2I prompt answers these questions, in roughly this order:
- Subject — who or what is in the frame, with concrete details (not "a person" but "a weather-beaten fisherman in a canvas jacket").
- Details — the textures and materials that make it specific: fabric weave, fur, skin, patina on metal.
- Scene & time — where it is and the atmosphere: an alley, a studio, a coastline. Time of day and weather set the palette more than any style word.
- Composition — how the frame is organized: rule of thirds, centered, leading lines, negative space, frame within a frame.
- Lighting — the look of the light: golden hour, rim light, volumetric god rays, neon, candlelight, studio soft light, chiaroscuro.
- Color — the palette direction: teal & orange, muted earth tones, pastel, monochrome, jewel tones.
- Style, medium & mood — the visual language (photorealistic, cinematic, anime, oil painting, pixel art) and the emotional tone (dramatic, peaceful, melancholic).
- Camera & format — shot size (close-up, full body, ultra-wide), focal length (35mm, 85mm), angle, and the aspect ratio you actually need.
Worked example
Assembled from the eight parts above, for a photorealistic portrait:
"a weather-beaten fisherman with a sun-roughened face, salt-stained canvas jacket, standing at the end of a small pier at dawn, mist rising off the water, centered composition with negative space above, soft golden-hour light from the left, muted earth tones, photorealistic photograph, melancholic mood, 85mm portrait lens, medium close-up, 4:5"
Count it: subject, texture, place, time, mist, composition, light, color, style, medium, mood, focal length, shot size, format — fourteen elements, each clause doing one job, none of it decoration. A short quality phrase at the end — "ultra-detailed, natural lighting" — reads as a quality target for the model.
Visual vocabulary image models understand
Image models are trained on photos and artworks annotated with technical terminology, so the standardized words work. The ones you will use most:
- Shot size — extreme close-up, close-up, medium shot, full shot, long shot, extreme wide, overhead, over-the-shoulder, two-shot.
- Angle — eye level, low angle, high angle, bird's-eye, worm's-eye, dutch angle, point of view (POV).
- Focal length & depth — 24mm wide, 35mm street, 50mm, 85mm portrait, 135mm telephoto, macro 1:1; shallow depth of field with creamy f/1.4 bokeh, or f/16 for everything sharp.
- Lighting — golden hour, blue hour, backlit, rim lighting, volumetric light, god rays, neon-lit, candlelit, moonlit, high-key, low-key, Rembrandt, chiaroscuro, dramatic shadow.
- Film & texture — 35mm film photograph, Kodak Portra 400, grainy film texture, light leak, direct flash from a point-and-shoot, VHS still.
- Art medium — oil painting impasto, watercolor wash, charcoal sketch, ink line art, colored pencil, pixel art, isometric art, cel-shaded anime, comic halftone, pop art, ukiyo-e woodblock, glitch art.
You do not need to name all of them. One or two technical choices — "85mm portrait lens, shallow depth of field" or "oil painting impasto" — often move the result more than three style adjectives.
Matching the prompt to the model
The major image models come from different labs: Midjourney V7 (Midjourney), DALL·E 3 (OpenAI), Flux.1 (Black Forest Labs), Imagen 3 (Google) and Stable Diffusion 3.5 (Stability AI), plus the Chinese models Seedream (ByteDance), Kolors (Kuaishou), Wan and Z-Image Turbo (Alibaba), MiniMax Image and ERNIE ViLG (Baidu). Their prompting strategies differ in two practical ways:
- Language. The Chinese models are native in Chinese — Chinese prompts read most naturally for them; the rest were trained mainly on English. Write the prompt in the language the model is strongest in.
- Quality phrase. Each model has a preferred finishing phrase — Flux.1 "photorealistic, accurate detail", Midjourney V7 "cinematic, highly detailed, artistic composition". Appending the right one makes the same selection read as a native prompt for each model.
The AI Image Prompt Generator handles both: it matches the output language to the model (Chinese for the Chinese models on the Chinese page, English everywhere else) and appends each model's own quality suffix.
Seven mistakes that flatten AI images
- Vague subject. "A person somewhere" leaves the model inventing the cast; be concrete about who or what is in the frame.
- Conflicting styles. "Photorealistic oil painting" cancels itself; pick one visual language and commit.
- No lighting intent. Without a light direction the model defaults to flat, neutral, catalog-style light; name your light.
- Adjective soup. Twenty style words dilute every instruction; ten concrete descriptors usually beat fifty adjectives.
- Mixed languages. Keep the prompt in one language; half-translated prompts confuse the model.
- Ignoring aspect ratio. The square default is wrong for wallpapers, banners, and book covers — declare 16:9, 4:5, or 3:2 up front.
- Negatives instead of positives. "No blur, no extra fingers" works less reliably than describing what you want: sharp focus, clean hands, one subject.
Build prompts with structure
Instead of free text, you can pick options from a structured grid. The AI Image Prompt Generator covers ten categories — subject type, scene, composition, lighting, color palette, art style, medium, mood, camera, and aspect ratio — with a reference image for every option, and assembles your single choice per category into one prompt tuned to the target model. It runs entirely in the browser: nothing you write is uploaded.