Prompt Engineering

The Complete AI Video Prompt Guide (2026 Edition)

The seven-part framework we use to write AI video prompts that behave predictably, plus tested examples for cinematic, ultra realistic, anime, commercial, social, and YouTube Shorts work.

R
ReelVision Editorial
Editorial Team
August 5, 2026 31 min read
A film director's notebook open on a matte black desk beside a clapperboard and cinema lens under warm gold light

Most people who are disappointed by AI video are not using a weak model. They are handing a capable model an ambiguous brief. A video model has to decide, in one pass, what the subject looks like, how it moves, where the camera sits, how the light falls, how long each beat lasts, and what the scene sounds like. Every decision you leave open is a decision the model makes for you, and it makes those decisions from the average of everything it has seen. Average is exactly what makes a clip look generic.

This guide is the reference we use internally at ReelVision AI when we write, test, and repair prompts. It covers the structure of a prompt that behaves predictably, the vocabulary that models actually respond to, and complete worked examples for cinematic, ultra realistic, anime, commercial, social, and YouTube Shorts work. Nothing here is theoretical. Every pattern below has been run through production video engines, compared against alternatives, and kept because it changed the output.

What a video model actually reads

A text to video model does not read your prompt the way a human reads a screenplay. It converts your words into a dense numeric description of a scene, then generates frames that stay consistent with that description over time. Two consequences follow, and almost every prompting rule in this guide comes from them.

First, concrete nouns and physically observable descriptions carry more weight than abstract quality words. The phrase beautiful cinematic masterpiece contains no visual instruction. The phrase low sun raking across wet asphalt, long shadows, haze in the air contains four. The second phrase moves pixels. The first one mostly wastes attention.

Second, the model has a limited attention budget per generation, and it spends more of that budget on whatever appears early and repeatedly. If your camera instruction sits in the last clause of a hundred word paragraph, it competes with everything before it. If a detail matters, it should be stated early, once, in plain language, and not contradicted later.

There is a third practical reality worth naming. Video models are strongest at short, physically plausible motion and weakest at anything requiring long chains of cause and effect. A person pouring coffee and glancing at a window is well inside the envelope. A person walking into a room, arguing with someone offscreen, then storming out and slamming a door is three shots pretending to be one. Splitting that into three prompts will beat any amount of clever wording in a single prompt.

Weak instructionWhy it failsStrong instruction
Beautiful cinematic shot of a womanNo subject detail, no camera, no light, no actionWoman in her thirties, dark wool coat, standing still on a train platform, shot at eye level on a 50mm lens
Epic dramatic lightingAbstract adjective with no light sourceSingle hard light from screen left, deep falloff, practical fluorescent flicker overhead
Amazing camera workThe model has to invent the moveSlow dolly in, roughly one meter over the shot, ending on a medium close up
4k, 8k, ultra HD, masterpieceResolution is a render setting, not a descriptionShallow depth of field, fine skin texture visible, natural motion blur on the hands
The same idea, written two ways

The seven-part prompt framework

Every reliable prompt we write contains the same seven components in roughly the same order. The order matters because it front loads the parts the model cannot guess and leaves the parts it can infer for last. You do not need all seven in every prompt, but you should know which one you are deliberately leaving out.

  1. Subject: who or what is on screen, described physically
  2. Action: the single motion that happens during the clip
  3. Environment: where it happens, and the surfaces and weather that define it
  4. Camera: position, lens feel, and one move
  5. Light: source, hardness, direction, and color temperature
  6. Sound: what should be audible if the engine generates native audio
  7. Constraints: aspect ratio, duration, and what must not appear

1. Subject: describe the body, not the vibe

Subject drift is the most common failure in AI video, and it usually starts with a vague subject line. A prompt that says a young man gives the model permission to reinvent that man on every generation. A prompt that says a man in his late twenties, close cropped black hair, light stubble, olive skin, wearing a grey crewneck sweater constrains hundreds of pixels of face and clothing before the first frame is drawn.

Describe subjects in this order: age range, hair, skin, distinctive feature, clothing, then anything held or worn. Keep the total under about twenty five words for one subject. Two subjects in a frame need clear separation, for example the taller of the two on the left, so the model can assign details correctly instead of blending them into one person. If you plan to reuse a person across shots, lock the description word for word and treat it as a fixed string you paste into every prompt. We go much deeper on this in Character Consistency: How to Keep the Same Face Across Every Shot.

2. Action: one verb, one beat

A single clip should contain one action with a clear beginning and end. Write the action as a present tense verb phrase describing physical movement, then add a small secondary motion if the clip is longer than six seconds. A good action line reads like stage direction: she sets the mug down and turns her head toward the window. A bad action line reads like a plot summary: she realizes she has been betrayed and decides to leave.

Motion vocabulary is worth being precise about. Slowly, gently, and abruptly change output measurably because they map to speeds the model has seen labeled that way. Emotional adverbs like nervously produce inconsistent results because nervousness looks different in every reference the model has learned. Prefer the physical tell over the emotion: her fingers tap once against the ceramic reads as nerves without asking the model to interpret an abstraction.

3. Environment: surfaces, depth, and weather

Environments read as real when they contain materials that respond to light. Wet asphalt, brushed steel, dusty glass, worn plaster, and steam all give the renderer something to do. Empty rooms and clean gradients give it nothing, which is why so many AI interiors look like showrooms. Name two or three surfaces and one atmospheric element and you have solved most of the realism problem before you touch the camera.

Depth cues matter as much as surfaces. Adding foreground occlusion, for example a doorframe edge in the near left of frame, or background separation, for example distant traffic lights out of focus, forces the model to build a scene with layers instead of a flat plate. Layers are what make a frame feel photographed.

4. Camera: position first, move second

State where the camera is before you state what it does. Eye level, low angle from waist height, high three quarter angle, and over the shoulder all set framing instantly. Then pick exactly one move: static, slow push in, slow pull out, lateral track, handheld follow, or crane down. Stacking two moves in one clip is the fastest way to get warped geometry, because the model has to reconcile two different motion fields across the same pixels.

Lens language works, within limits. Wide angle, 35mm, 50mm, 85mm, and macro all shift perspective and depth of field in the expected direction. Anamorphic gives you oval highlights and horizontal flare. What does not work is invented gear specificity: naming an exact camera body and firmware rarely changes anything, and it eats attention that a light or surface detail would have used better. The full vocabulary, with the moves that survive contact with real engines, is in Camera Language: The Moves That Make AI Footage Feel Directed.

5. Light: name the source

Amateur prompts describe the mood of the light. Professional prompts name the source and let mood follow. Single bare bulb overhead, late afternoon sun through venetian blinds, sodium street lamp behind the subject, and monitor glow as the only source each produce a distinct and repeatable look. Add hardness with hard or diffused, add direction with from screen left or backlit, and add temperature with warm tungsten or cool overcast daylight.

One more control is underused: contrast and falloff. Deep shadows with no fill reads as drama. Soft even fill with lifted blacks reads as documentary or commercial. Saying which you want is faster and more reliable than listing film stocks and hoping the model recovers your intent from them.

6. Sound: prompt it if the engine generates it

Modern engines can generate native audio in the same pass as the picture, which means the audio is locked to the motion instead of being layered on afterward. If you say nothing about sound, you get whatever ambience the model associates with the scene, and it is often close to silence. Two or three specific cues change that: rain on a metal awning, distant traffic, no music, or single male voice, calm, close to the microphone.

Be explicit when you do not want something. No music is a real instruction that works. So is no dialogue. Music in particular tends to appear uninvited and can make an otherwise usable clip unusable in an edit, because you cannot separate it from the effects track afterward.

7. Constraints: format, duration, exclusions

Constraints are the least interesting part of a prompt and the most likely to save a render. State the aspect ratio in words when the platform expects it, keep duration inside what the engine natively supports rather than asking for a length it will silently truncate, and list a short exclusion set for the artifacts your subject invites. A crowd scene invites duplicated faces. A vehicle scene invites melted body panels. A text heavy scene invites garbled lettering, which is why we recommend never asking a video model for readable words on screen and adding titles in your editor instead.

The framework as one prompt

A woman in her late thirties, shoulder length auburn hair, freckled skin, wearing a faded denim jacket, stands at a kitchen counter and slowly pours coffee into a white ceramic mug, then lifts her eyes toward the window. Small apartment kitchen, chipped tile backsplash, condensation on the glass, steam rising from the mug. Camera at eye level on a 50mm lens, static, framed as a medium shot with the doorframe edge softly occluding the near left of frame. Late morning sun through the window as the only source, diffused, cool daylight, gentle falloff into the shadowed room. Audio: pouring liquid, faint street noise outside, no music. Vertical 9:16, natural motion, no on screen text.

Every clause maps to one of the seven components. You can delete any single clause and predict exactly what the model will start guessing.

Prompt length, order, and the attention budget

There is a sweet spot for prompt length, and in 2026 it sits between roughly 60 and 120 words for a single clip. Below 60 words you are usually leaving light, camera, or environment unspecified, which means the model fills those in from its own averages. Above about 150 words, later clauses start to lose influence, and contradictions become almost inevitable because a long prompt is hard to keep internally consistent.

The order that works, tested across engines, is subject and action first, environment second, camera third, light fourth, audio fifth, constraints last. There is a good reason for that order beyond habit. Subject and action define what the frames are of. Environment and camera define how it is seen. Light and audio are modifiers on top. Constraints are metadata. When you reverse this and lead with lens and film stock talk, the model spends its first and strongest attention on aesthetics and gets to your subject already committed to a look.

  • Say each thing once. Repetition for emphasis is a habit from image models and it hurts video coherence.
  • Prefer three specific details over ten generic ones. Ten adjectives dilute each other.
  • Delete every word that does not change a pixel or a sample. Stunning, breathtaking, and award winning change neither.
  • Keep numbers plausible. Asking for a 200mm macro handheld drone shot describes a rig that cannot exist, so the model picks one interpretation at random.

Cinematic prompts: how to earn the film look

The film look is not a filter. It is the combination of controlled light, a committed lens choice, deliberate camera height, and restrained motion. When people say a clip looks cinematic, they usually mean four specific things are present: separation between subject and background, a light source you could point at, motion that has weight, and framing that leaves negative space on purpose.

Write cinematic prompts around a single dramatic idea per shot. The most common mistake is trying to make one clip carry an entire scene, which produces mush. Instead, write the establishing shot, the medium, and the close as three prompts that share an environment description word for word. Reusing the environment string is what makes them cut together, and it costs nothing.

Cinematic, night exterior

A man in his fifties, grey stubble, heavy canvas work jacket, stands motionless in the middle of an empty two lane road and looks back over his shoulder. Rural night, wet asphalt reflecting a single sodium street lamp, tall grass on both shoulders, thin ground mist. Camera low, near ground level, static, 35mm lens, subject centered and small in frame. The street lamp is the only light source, hard and warm, strong falloff into black at the frame edges. Audio: crickets, distant wind, no music. Cinematic 16:9, no on screen text.

One subject, one action, one light, one lens. The drama comes from framing him small in a wide, dark frame.

Cinematic, interior dialogue coverage

Close up of a woman in her thirties, dark hair pulled back, wearing a charcoal blazer, seated at a wooden table. She listens, then lowers her gaze. Dim restaurant interior, brass fixtures, half empty wine glass in soft foreground blur. Camera at eye level, 85mm lens, very slow push in over the shot. Warm practical light from a table candle from screen left, diffused fill from a window behind camera, deep shadow on the right side of her face. Audio: low room murmur, cutlery, no music. Cinematic 16:9.

Slow push in on a long lens is the safest way to add motion to a dialogue clip without breaking the face.

For period, genre, or franchise looks, describe the physical production design rather than naming a film. Grain, halation, and slightly desaturated color with lifted blacks gets you a seventies feel. Naming a specific movie usually returns a vague approximation and sometimes returns nothing at all. If you want a full breakdown of shot construction for dramatic work, The Anatomy of a Cinematic AI Video Prompt walks through one shot clause by clause, and the AI movie scene generator is set up for exactly this kind of coverage.

Ultra realistic prompts: the documentary discipline

Realism and cinematic polish pull in opposite directions. Cinematic wants controlled light and elegant motion. Realistic wants imperfect light and imperfect operating. If you want footage that reads as something a person actually shot, you have to prompt in flaws on purpose.

The single most effective realism technique is naming the capture device. Handheld smartphone footage, dashcam mounted behind the windshield, action camera on a chest mount, and consumer camcorder each carry a bundle of expectations the model already knows: rolling shutter, auto exposure hunting, limited dynamic range, compression, and a specific kind of shake. Naming the device gets you all of it in three words, which is far more efficient than describing each artifact.

  • Add one imperfection: slight auto exposure shift, minor lens smudge, or a brief reframe as the operator corrects.
  • Keep the light unflattering. Overcast midday, fluorescent office ceiling, and mixed streetlight all read as real.
  • Avoid perfect symmetry and avoid slow motion. Both signal a production.
  • Skip the word cinematic entirely. It pulls the render toward polish and away from believability.
Ultra realistic, handheld street

Handheld smartphone footage, vertical, of a delivery courier in a yellow rain shell wheeling a bicycle along a crowded sidewalk. Rain has just stopped, puddles on uneven paving, storefront windows fogged, pedestrians passing in both directions. Camera roughly chest height, following behind at walking pace, natural hand shake, one small reframe as the operator adjusts. Overcast daylight, flat and even, no strong shadows. Audio: footsteps, wet tires, muffled traffic, no music. 9:16, natural motion blur, no on screen text.

Nothing here is beautiful, which is the point. Flat light plus device naming plus one operator error.

Ultra realistic, vehicle interior

Dashcam footage mounted behind the windshield of a car driving a two lane road at dusk. Oncoming headlights, insects on the glass, wipers at intermittent speed clearing light drizzle. Fixed mounted camera, wide angle, slight vibration from the road surface, mild lens distortion at the frame edges. Fading daylight with headlights and tail lights as the practical sources, limited dynamic range with clipped highlights. Audio: engine hum, wiper sweep, tire noise on wet road, no music. 16:9, no on screen text, road surface matte and not mirror like, vehicle bodies solid with consistent panels.

The last clause is a targeted exclusion. Vehicles and wet roads are where realism prompts most often break.

Realistic work benefits from the engine as much as the wording. Our realistic AI video generator routes to models tuned for physical plausibility and native audio rather than stylized rendering, and it applies capture aware constraints before the prompt reaches the provider.

Anime and stylized prompts

Anime prompting inverts one of the rules above. Where realistic work needs physical light sources, stylized work needs a described rendering technique, because the look is drawn rather than photographed. You are describing an artistic process, so name line weight, shading method, and color treatment.

Useful vocabulary: cel shaded with hard shadow edges, thin consistent line art, flat color fills with limited gradients, painted background with visible brush texture, speed lines during fast motion, and rim light on hair. Camera language still applies, but it should be simpler. Animated sequences favor held frames with a single push or pan rather than complex dolly work, because that is how the reference material was made.

Anime, action beat

Cel shaded anime style. A teenage swordfighter with spiked white hair, dark uniform coat, red scarf trailing behind, dashes across a rooftop and leaps toward the frame. Night city rooftops, painted background with visible brush texture, neon signs glowing below, laundry lines swaying. Camera at low angle, single fast pan following the leap, held on impact. Hard shadow edges, cool blue ambient with warm rim light on hair from the signs below, flat color fills. Speed lines during the dash. Audio: wind, fabric snap, footsteps on metal, no dialogue. 16:9, no on screen text.

Style declared first, then the standard framework. Style first is the one case where leading with aesthetics is correct.

Stylized, quiet character moment

Soft anime style with thin line art and watercolor backgrounds. A girl in a school uniform sits on a train, chin resting on her hand, watching the window as she blinks slowly. Late afternoon train interior, worn fabric seats, dust in the light, blurred green fields passing outside. Camera static, medium close up at eye level, slight parallax in the window only. Warm low sun from the window at screen right, gentle gradient shading, muted palette. Audio: rhythmic track noise, faint announcement chime, no music. 16:9.

Restricting the motion to the window parallax keeps the character stable, which is where stylized clips usually fail.

Commercial and product prompts

Commercial work has a different success test. It is not about drama or believability, it is about whether the product reads clearly, looks desirable, and holds the frame. That means the product needs to be the brightest and sharpest thing in the shot, the motion should be short and clean, and the surfaces need to behave correctly, especially glass, liquid, and brushed metal.

The reliable commercial pattern is a locked or slowly moving camera, a controlled studio light plan, and one physical event: a pour, a drop, a lid closing, steam rising, a hand entering frame. Hands are risky in AI video, so keep them partially out of frame or moving slowly across a short distance. Never ask for a logo or label text to be legible. Shoot the product clean and add branding in your editor, where it will be sharp and correct.

Commercial, beverage hero

A frosted glass bottle of sparkling water stands on a wet slate slab as condensation runs down the side and a single drop falls onto the stone. Dark studio background, faint mist drifting behind, water pooling on the slate. Camera at product level, macro feel, very slow lateral track to the right, shallow depth of field. Two light setup: hard rim light from behind screen left creating a bright edge on the glass, soft large source from the front right for gentle fill, deep background falloff. Audio: fizz, a single water drop, quiet room tone, no music. 16:9, no visible text or labels, no hands.

Rim light plus a dark field is the standard beverage plan, and it works on video models exactly as it does on set.

Commercial, lifestyle vertical

A pair of hands lifts a matte black wireless earbud case from a linen surface and opens the lid in one smooth motion. Bright minimal interior, linen texture, a ceramic cup and a folded newspaper softly out of focus behind. Camera looking down at a slight angle, static, 50mm feel, tight framing on the hands and product. Large soft window light from screen left, subtle shadow under the case, clean neutral white balance. Audio: a soft click as the lid opens, quiet room tone, no music. 9:16, no visible branding or text, fingers with correct anatomy and slow deliberate movement.

Slow deliberate movement is a real instruction for hands, and it measurably reduces finger artifacts.

Social media prompts: built for the first second

Social video is judged in the first second, in a vertical frame, at a small size, often with sound off. That changes what a good prompt looks like. Composition has to be tighter, because detail is lost at phone size. The most interesting thing has to already be happening at frame one, because there is no time for a reveal. And the top and bottom of the frame belong to the platform interface, so subjects need to sit in the middle band.

  • Open on motion, not on a static establishing frame. State the action as already in progress.
  • Keep the subject in the central two thirds of a 9:16 frame so captions and platform UI do not cover it.
  • Choose high contrast between subject and background so the clip still reads at thumbnail scale.
  • Design for silence first, then let the native audio be a bonus rather than the whole idea.
Social, scroll stopping opener

A barista pulls a shot of espresso as thick crema streams into a clear glass, already in motion at the first frame. Small specialty cafe, brushed steel machine, steam wand hissing, warm wood counter, blurred customers behind. Camera close and slightly low, static, tight on the portafilter and glass, shallow depth of field. Warm overhead pendant light from above screen right, hard highlight on the steel, dark background separation. Audio: espresso machine hum, liquid stream, no music. Vertical 9:16, subject centered in the middle band of frame, no on screen text.

The subject is already moving, the contrast is high, and everything important sits away from the frame edges.

There is a lot more to the vertical format than reframing a horizontal idea, including how to structure a three beat vertical story and where to put the cut. We wrote that up in Vertical AI Video That Actually Holds Attention and it pairs directly with this section.

YouTube Shorts prompts: sequences, not single clips

Shorts reward a complete small story inside a fixed vertical frame, which means you are almost never writing one prompt. You are writing four to eight prompts that share a visual world, then cutting them together. The prompt craft question becomes continuity: how do you keep the subject, location, and light consistent across separate generations.

The technique is a shared block and a variable block. Write your subject description and your environment description once, treat them as fixed text, and change only the action and camera lines per shot. This is unglamorous and it is the single highest leverage habit in multi shot AI video.

ShotAction lineCamera line
1walks into the workshop and flicks on the overhead lightwide, static, low angle from the doorway
2sets a warped guitar neck on the bench and runs a thumb along itmedium close up, static, 50mm feel
3clamps the neck and begins sanding in short even strokestight on the hands, slow lateral track
4blows dust off the surface and holds it up to the window lightmedium, slow push in
5steps back and turns the finished neck slowly in both handswide, static, matching shot one framing
A five shot Shorts sequence built from one shared block
Shorts, shared block plus shot three

SHARED: A man in his forties, close cropped greying hair, wire rimmed glasses, wearing a faded brown apron over a grey shirt. Small woodworking workshop, sawdust on the floor, hand tools on a pegboard, one dusty window at screen right, single warm overhead bulb. Audio: room tone, tool contact, no music. Vertical 9:16, no on screen text. SHOT 3: He clamps a guitar neck to the bench and begins sanding in short even strokes, dust catching the window light. Camera tight on his hands, slow lateral track to the left, shallow depth of field.

Paste the shared block into all five prompts unchanged. Only the shot line moves. Continuity across the cut comes for free.

For longer form pieces built the same way, the AI short film generator and the AI trailer generator are structured around multi shot sequences rather than single clips, which saves you assembling the shared block by hand.

Multi shot sequences and continuity

Continuity in AI video is a prompt discipline, not a post production fix. Four things have to stay stable across shots: the subject, the wardrobe, the location, and the light direction. If the light moves from screen left to screen right between two clips, the cut will feel wrong even if the viewer cannot say why, because that is the same error that breaks continuity in live action.

The workflow we recommend is a continuity sheet, which is a plain text file with three locked strings and one list. The locked strings are subject, environment, and light. The list is your shot order with an action line and a camera line each. Generate shot by shot, and when a shot drifts, regenerate that shot only rather than rewriting the sheet. Rewriting the sheet mid sequence is how people lose an afternoon.

  1. Lock the subject string and never edit it once a shot has passed review.
  2. Lock the environment string, including surfaces and time of day.
  3. Lock the light direction, stated relative to the camera as screen left or screen right.
  4. Vary only action and camera per shot, and change one of the two at a time when troubleshooting.
  5. Keep a rejected list so you do not repeat a wording that already failed.

Image to video prompting

When you supply a starting image, your prompt stops being a description and becomes a set of instructions about change. This is the most common prompting mistake in image to video work: people paste their text to video prompt in, the model receives a description of things it can already see, and the redundancy competes with the motion instruction.

Write image to video prompts as motion only. Name what moves, how fast, and what the camera does. Mention the subject only enough to attach the motion to the right thing, for example the woman on the left turns her head slowly. Do not describe the light or the environment at all unless you want them to change, because restating them invites the model to reinterpret them.

AspectText to videoImage to video
Subject detailFull physical description requiredOne identifying phrase only
EnvironmentTwo or three surfaces plus atmosphereOmit unless it should change
LightSource, hardness, direction, temperatureOmit unless it should change
CameraPosition and one moveOne move only, position is already fixed
ActionOne beat, present tenseOne beat, described as a change from the still
Typical length60 to 120 words15 to 40 words
Text to video versus image to video for the same shot
Image to video, correct form

The woman lowers her cup and turns her head slowly toward the window. Steam continues to rise. Camera pushes in very slightly. Everything else stays as it is.

Short, motion only, no restatement of the frame. Longer is worse here, which surprises almost everyone the first time.

Image conditioning is also the right tool for restoration and continuation work on real footage rather than generated frames. If you are working from archival material, the constraints are different again, and Restoring Old Family Footage Without Making It Look Fake covers the parts that are specific to real source media.

Audio prompting for native sound models

Native audio generation changed prompt structure more than most people have adjusted for. When picture and sound come from the same pass, the sound is synchronized to what actually happened in the frame, so footsteps land on footfalls and a door latch clicks when the latch closes. That is worth prompting for deliberately, because the default is usually a thin room tone.

Structure an audio clause in three layers. Ambience is the bed, for example rain on glass or office ventilation. Effects are the events, for example a chair scraping or a shutter click. Voice is optional and should specify one speaker, a tone, and a distance, for example single female voice, low and unhurried, close to the microphone. Then state your exclusions, because no music and no dialogue are the two most useful audio instructions in the whole guide.

  • Three cues maximum. More than three and the mix turns to soup.
  • Name distance for voices, since close to the microphone and heard from across the room sound completely different.
  • Ask for no music by default. Add it in your edit where you can duck and time it.
  • If you need precise narration wording, record or synthesize it separately and use the native audio only for ambience and effects.

Negative prompts and what never to write

Negative prompting works best when it is short and targeted at the specific artifact your scene invites. A long generic exclusion list copied between projects mostly wastes attention. Think about what your subject makes likely: crowds invite duplicated faces, hands invite extra fingers, vehicles invite deformed panels, water invites mirror surfaces, and anything with signage invites garbled lettering.

Scene containsLikely artifactExclusion to add
Hands doing fine workExtra or fused fingerscorrect hand anatomy, slow deliberate movement
Vehicles in motionMelted panels, floating wheelssolid vehicle bodies, consistent panels, wheels in contact with the road
Wet streets or waterMirror like unnatural reflectionsmatte wet surface, diffuse reflections
Crowds or background peopleDuplicated or morphing facesbackground faces soft and out of focus
Signs, screens, packagingGarbled invented letteringno on screen text, no readable signage
Animals runningLeg count and gait errorscorrect four legged gait, legs in contact with the ground
Targeted exclusions by subject type

There is also a short list of things that belong in no prompt. Resolution tokens like 4k and 8k do nothing, because output resolution is a render setting. Quality tokens like masterpiece and award winning are inherited from image model habits and add no visual information. Naming a real actor or a copyrighted character is both a compliance problem and a quality problem, since the result is usually a poor approximation. And asking for readable text on screen will fail on every current engine, which is why titles belong in your editor.

Debugging a prompt that failed

When a clip comes back wrong, resist the urge to rewrite the whole prompt. Rewriting changes ten variables at once and teaches you nothing. Diagnose the failure category first, then change one clause. Almost every failure falls into one of five buckets, and each bucket has a specific fix.

SymptomReal causeThe one change to make
Subject looks different than describedSubject clause too short or buriedMove the subject clause first and add two physical details
Warped geometry during motionTwo camera moves, or a move that is too fastReduce to one move and add the word slow
Nothing much happensAction written as intention instead of movementReplace with a present tense physical verb phrase
Flat, plasticky lookNo named light source and no surface materialsName one light source and two surfaces
Unusable audioNo audio clause, so the model improvisedAdd ambience plus one effect plus no music
Five failure categories and the single clause that fixes each

Keep a log. Two columns is enough: the prompt you ran and one sentence about what happened. Within twenty generations you will have a personal reference of what your preferred engine responds to, which is worth more than any published prompt list, including this one. Prompting skill is mostly accumulated evidence about a specific model.

A repeatable twenty minute workflow

Here is how the whole guide compresses into a working session. This is deliberately dull, because reliability comes from a boring process rather than inspiration.

  1. Write the one sentence purpose of the video. If you cannot, do not generate yet.
  2. Write the shot list as action lines only, one line per clip, before writing any full prompts.
  3. Write the three locked strings: subject, environment, light.
  4. Assemble shot one from the locked strings plus its action and camera line, and generate it.
  5. Judge shot one on subject, motion, light, and audio separately, and fix only the weakest of the four.
  6. Once shot one is right, generate the rest without changing the locked strings.
  7. Regenerate individual failures rather than the sequence.
  8. Assemble in your editor and add titles, music, and any precise narration there.

Prompt library: six starting points

These are compressed versions of the patterns above, written to be edited rather than used verbatim. Swap the subject and environment, keep the camera, light, audio, and constraint structure, and you will keep most of the reliability.

1. Cinematic establishing shot

[SUBJECT] stands at [LOCATION] and [SINGLE ACTION]. [TWO SURFACES], [ONE ATMOSPHERIC ELEMENT]. Camera [HEIGHT], static, 35mm lens, subject small in a wide frame. [ONE LIGHT SOURCE] as the only source, hard, from screen left, deep falloff. Audio: [AMBIENCE], [ONE EFFECT], no music. 16:9, no on screen text.

2. Ultra realistic handheld

Handheld smartphone footage of [SUBJECT] [SINGLE ACTION]. [LOCATION], [TWO SURFACES], ordinary clutter. Camera chest height, following at walking pace, natural hand shake, one small reframe. Overcast daylight, flat and even. Audio: [AMBIENCE], footsteps, no music. 9:16, natural motion blur, no on screen text.

3. Anime action beat

Cel shaded anime style. [SUBJECT] [FAST ACTION]. [LOCATION], painted background with brush texture. Camera low angle, single fast pan following the action, held on impact. Hard shadow edges, cool ambient with warm rim light on hair, flat color fills, speed lines. Audio: wind, fabric, impact, no dialogue. 16:9, no on screen text.

4. Product hero

[PRODUCT] on [SURFACE] as [ONE PHYSICAL EVENT]. Dark studio background, faint mist. Camera at product level, macro feel, very slow lateral track, shallow depth of field. Hard rim light from behind screen left, soft fill from front right, deep background falloff. Audio: [ONE EFFECT], quiet room tone, no music. 16:9, no visible text or labels, no hands.

5. Vertical social opener

[SUBJECT] [ACTION ALREADY IN PROGRESS AT FRAME ONE]. [LOCATION], [TWO SURFACES], blurred background activity. Camera close and slightly low, static, tight framing, shallow depth of field. [ONE WARM PRACTICAL LIGHT] from above screen right, hard highlight, dark background separation. Audio: [AMBIENCE], [ONE EFFECT], no music. 9:16, subject centered in the middle band, no on screen text.

6. Shorts shared block

SHARED: [SUBJECT, 20 TO 25 WORDS]. [ENVIRONMENT, SURFACES, TIME OF DAY, ONE LIGHT SOURCE AND ITS DIRECTION]. Audio: [AMBIENCE], no music. Vertical 9:16, no on screen text. SHOT [N]: [ACTION LINE]. Camera [FRAMING], [ONE MOVE].

What changed in 2026, and what did not

Three things genuinely changed this year. Native audio became standard rather than a novelty, which means audio clauses now belong in every prompt. Image conditioning got strong enough that continuing a shot from its own last frame is the default answer to continuity, not a workaround. And physical plausibility improved to the point that flat, unflattering light is now a realism tool instead of a liability, because engines can render it without falling apart.

What did not change is the part people keep hoping will. Models still cannot render reliable on screen text. They still cannot carry a multi beat narrative inside one clip. They still reward concrete physical description over aesthetic adjectives, and they still punish contradictions. Every one of those limits has been stable across three generations of engines, which is a good reason to build your workflow around them rather than against them.

If you want a broader picture of how the pieces fit together before you start, What is ReelVision AI explains the pipeline, and Disaster Scenes That Feel Real Instead of Rendered is a good example of applying these rules to a genuinely hard subject with the AI disaster scene generator.

Put this guide to work

Copy any prompt above, change the subject, and generate it. The framework holds up better under editing than it does under reading.

Start generating
Share

Frequently asked

R
Written by
ReelVision Editorial
Editorial Team

The ReelVision AI editorial team writes about cinematic AI video production, prompt craft, and the workflows behind believable footage.

Try this in ReelVision AI

Everything in this article is one prompt away. Open the studio and render it with locked characters, native audio, and real cinematic motion.

Start Creating
Newsletter

Cinematic AI craft, weekly

One email. Prompt breakdowns, camera language, and workflows that make AI footage look directed.

No spam. Unsubscribe in one click.