Tutorial

Cinematic AI video prompting: direct the model like a cinematographer.

Exploded cinema lens with a blue light path through its elements

Cinematic AI video prompting means describing a shot the way a cinematographer would: subject and action, shot size, lens, camera movement, lighting direction, and palette, in that order. It does not mean adding the word “cinematic.” A model has no camera to infer a framing from a mood word and no gaffer to infer a light source from an adjective. Give it the same decisions a director gives a crew and the output reads as directed, not decorated.

The difference shows up in the footage. Creatives who get controllable results out of Runway or Kling learned the vocabulary of a film set. Creatives who spend an afternoon re-rolling the same prompt learned a mood board.

The short version
  • Cinematic results come from cinematographic vocabulary, not adjective piles. “Cinematic, 8k, masterpiece” gives the model nothing to act on.
  • Build the prompt in order: subject and action, shot size, lens, movement, lighting, palette, mood last.
  • Change one variable per iteration. Contradictory or over-specified prompts fight the model instead of directing it.
  • Movement language that names no physical rig tends to come back as a zoom.

What makes an AI video prompt cinematic?

A cinematic prompt names the mechanism, not the effect. “A woman walks through a rain-lit alley, cinematic, moody, 8k” describes a feeling and hopes the model shares your taste. “A woman walks through a rain-lit alley, medium shot, 35mm lens, shallow depth of field, slow dolly-in, key light from a neon sign camera-left, cool blue palette” describes a set of choices, each one a term the model can map to a visual pattern. The glossary has entries for most of these terms if you need a lookup mid-session.

Text-to-video and image-to-video models learn from captioned footage, and caption vocabulary borrows from the same place a shot list does: size, lens, movement, lighting setup. Terms like “low-angle,” “35mm,” and “dolly-in” sit beside a narrow band of images. “Cinematic” sits beside everything, because it gets tacked onto prompts regardless of what the frame shows. A word that correlates with nothing specific moves the output nowhere specific.

The anatomy of a cinematic prompt

A prompt is a sentence built from blocks in a fixed order, the order a shot gets built in: what happens, how it is framed, what lens sees it, how the camera moves, how it is lit, how it is graded. Mood comes last because mood is the result of the other six decisions.

PROMPT ANATOMY, IN ORDER Subject / action "a courier runs, hood up" Shot size wide shot, medium shot, close-up Lens 35mm, 85mm, shallow depth of field Movement static, pan, slow dolly-in, handheld Light side-lit, backlit, motivated by neon Palette teal and amber, desaturated Mood tense, quiet dread, if you use it at all The first six blocks are decisions a model can render. Mood is what they add up to. "A courier runs through a rain-lit alley, hood up. Wide shot, 35mm, slow dolly-in, backlit by a neon sign, cool blue and amber palette."
The seven blocks, in the order a shot gets built. Shot size is highlighted because it decides how much the other blocks need to say. Mood is the output of the first six, not an input alongside them.

What order should a video prompt follow?

Lead with subject and action: who or what, doing what, in a single clear clause. Then shot size, because framing decides what detail the rest of the prompt needs to specify. Then lens and depth of field, which set how the background behaves. Then camera movement. Then lighting: direction, quality, and a motivated source if one exists in the scene. Then palette and grade. Mood goes last, if you use it at all, because by that point the concrete choices have already produced one.

A prompt opening with “moody, atmospheric, cinematic” spends its strongest position on words with no fixed visual referent, and pushes the load-bearing nouns to the end. Put them first. The AI filmmaking guide walks through the same ordering at the level of a full shot list.

Shot size sets what everything else has to say

Shot size is the highest-leverage term in the anatomy. An extreme close-up needs no background note; an extreme wide shot needs no facial expression. Name the size first and you stop over-specifying the rest.

SHOT SIZE, ECU TO EWS extreme close-up close-up medium shot full shot wide shot extreme wide shot
From ECU, where a detail fills the frame, to EWS, where a figure reads as a mark in the environment. Choosing the size first tells the model how much of the scene needs describing at all.

Why do adjectives like cinematic and 8k not work?

They are compliments, not instructions. “Cinematic” asks the model to reverse-engineer an aesthetic tradition from one word and guess which part you mean: framing, grade, pacing, or a specific era’s lens choices. “8k” names an output resolution the prompt does not set. “Masterpiece” asks for nothing describable at all.

Overuse compounds the problem. A term that appears in nearly every caption regardless of what the frame shows stops correlating with any visual outcome, so it stops moving the output. “Side-lit,” “dolly-in,” and “35mm” pair with a narrower, steadier set of visuals, so they pull harder. Swap the compliment for the decision and keep the sentence the same length.

How do I control camera movement in AI video?

Name one movement, using a term tied to a physical rig: static, pan, tilt, slow dolly-in, dolly-out, or handheld. Avoid vague verbs like “the camera moves toward” or “flies over,” which commonly come back as a zoom. A dolly-in that arrives as a zoom means the movement language was too loose to rule the zoom out.

Test movement terms one at a time rather than stacking “slow dolly-in with a subtle pan and a handheld feel” into one clause. Three camera behaviors have to resolve into one continuous motion, and they tend to average into a soft, undirected drift. Isolate the term, confirm what it produces, then add a second movement if the shot needs one.

One variable at a time: a worked example

The fastest way to learn a model’s vocabulary is to hold everything constant and change one block. Here is the same scene taken through three passes.

  1. Flat. “A detective in a rain-lit alley, cinematic, moody, 8k, masterpiece.” No shot size, no lens, no movement, no light direction. The model guesses every one of them and averages toward generic footage.
  2. Partially directed. “A detective walks through a rain-lit alley at night, medium shot, slow dolly-in.” Better. Shot size and movement are decisions now, but lighting and palette are still undefined, so the grade will drift shot to shot.
  3. Directed. “A detective walks through a rain-lit alley at night, medium shot, 35mm, shallow depth of field, slow dolly-in, backlit by a neon sign, cool blue and amber palette.” Every block filled. The mood reads as tense without the word “tense” appearing anywhere.

Between passes, change only what you are testing. If the movement in pass two looks wrong, hold the shot size and lens and try a different movement term, not three new terms at once. That is the discipline of giving one note at a time to a working cinematographer: “same shot, hold the movement, warm the light,” not a rewritten brief. The shift toward conversational, stateful revision in newer video tools rewards that habit.

The courier shot, both ways

The anatomy diagram ends on a courier sentence. Here it is generated, beside the bare prompt most people write for the same shot. Both clips came out of Veo 3.1 Fast at the same settings, six seconds, no audio, one attempt each, so the prompt is the only variable.

Prompt 1: nothing named
a courier runs through an alley, cinematic, 8k
Two content words and two compliments. The model filled every empty block itself: overcast daylight instead of night, wet ground but no rain, a full shot at eye level, flat grey, and blocking that runs the courier at the camera for two seconds before turning him round to run away. “Cinematic, 8k” named no framing, no lens and no light, so nothing on screen traces back to those two words.
Prompt 2: every block named
A courier runs through a rain-lit alley, hood up. Wide shot, 35mm lens, slow dolly-in, backlit by a neon sign, cool blue and amber palette.
Night and rain came from “rain-lit,” the hood from “hood up,” the size of the figure from “wide shot,” the sign behind the subject from “backlit by a neon sign,” the grade from the palette clause. “Slow dolly-in” came back as something else: two seconds in, the courier turns and runs at the camera, so the frame tightens through blocking, and both the wide shot and the neon backlight are gone by the four-second mark. The sign itself says nothing. Its lettering is gibberish, and no block in the anatomy fixes that.

Where directed prompts still fail

Three patterns account for most bad output once the anatomy is in place.

  • Contradictory instructions. “Static shot, slow dolly-in” or “overcast, harsh sunlight” asks the model to satisfy two mutually exclusive terms. It picks one, blends them into something muddy, or oscillates across the clip. Read the prompt back as instructions to a crew. If a crew could not follow it, the model cannot either.
  • Over-specification. Adjectives stacked across every block compete instead of compounding. Three palette words or two lighting directions force an average, and the average is flatter and less committed than one clear choice per block.
  • Movement words that map to zooms. The widest gap between what a prompt asks for and what a clip delivers. If no physical rig can perform the move you named, do not expect that rig’s motion.

Questions creatives ask

What makes an AI video prompt cinematic? It names the decisions a cinematographer names on set: shot size, lens, camera movement, lighting direction, palette. Mood words alone give the model nothing to render, because there is no camera to infer a framing from and no light source to infer a quality of light from.

What order should a video prompt follow? Subject and action, shot size, lens and depth of field, camera movement, lighting, palette and grade, mood last. That mirrors how a shot gets built: you block the action before you choose the frame, and you light before you grade. Opening with style words buries the concrete decisions at the end.

Why do adjectives like cinematic and 8k not work? They describe a quality you want, not a decision the model can act on. A model can point a camera, choose a focal length, or place a key light. It cannot look up what cinematic means and reverse-engineer the framing that would produce it. Overuse dilutes those words further, so they move the output less than low-angle, 35mm, or side-lit.

How do I control camera movement in AI video? Name one movement, using vocabulary tied to a physical rig: static, slow dolly-in, handheld, pan. Vague phrasing commonly comes back as a zoom, so the rig-specific term is what rules the zoom out. Test one movement word at a time to learn which term maps to which motion.

How many things should I change between prompt iterations? One. Swap the lens, or the light, or the movement, never all three. Changing a single term isolates what caused the shift and keeps the parts that worked, the same way a director gives one note between takes.

Learn this beside the people building it.

Membership is free. Masterclasses from industry leaders, hackathons where you finish something the same day, and mentor circles matched to what you want to learn. For engineers and creatives alike, across film, design, image, sound, and story.