Best AI Video Prompts (2026): 100 Copy-and-Paste Prompts That Actually Work
One hundred original, tested AI video prompts across 20 categories, each with a copy button and a short explanation of why it renders cleanly instead of muddy.

Most prompt collections you find online are padded lists of the same twenty ideas with the words shuffled around. This is not that. Below are 100 original AI video prompts, written and tested for the way 2026 video models actually behave, and every single one comes with a short explanation of why it renders cleanly. If you only want the text, copy it and go. If you want to get better at this permanently, read the explanations, because they teach the pattern rather than the sentence.
Everything here is built on the same seven-part structure covered in The Complete AI Video Prompt Guide (2026 Edition): subject, action, environment, camera, light, sound, constraints. If a prompt below feels long, that is deliberate. Video models do not reward brevity the way image models sometimes do. They reward specificity in the six places where the render can go wrong, and silence everywhere else.
What makes a prompt work in 2026
Three things changed between 2024 and 2026, and they explain why old prompt lists fail. First, models generate native audio, so a prompt that says nothing about sound gets whatever the model guesses, which is usually a generic music bed nobody asked for. Second, motion coherence improved enough that the failure mode moved from melting limbs to wrong physics, so weight, friction, and momentum now belong in the prompt. Third, prompt adherence got sharp enough that vague style words like cinematic actively hurt, because the model applies its own averaged interpretation instead of yours.
The practical consequence is that a strong 2026 prompt reads less like a wish and more like a shot list entry a camera operator could execute. It names one subject, one continuous action, a place with two concrete details, a lens and a single camera move, a key light with direction, an ambience bed with one or two specific sounds, and a short constraint line that closes the door on text overlays and duplicated limbs. That is the whole recipe, and it is the recipe behind all 100 prompts below.
| Element | Beginner prompt | Advanced prompt |
|---|---|---|
| Subject | a woman | a woman in her late thirties, wet wool coat, hair flattened by rain |
| Action | walking around dramatically | walks eight steps toward camera and stops |
| Environment | a city | narrow brick alley, standing puddles, a fire door propped open |
| Camera | cinematic shot | medium shot, 35mm, slow push in, eye level |
| Light | moody lighting | single sodium streetlight behind her, hard rim, deep shadow on her face |
| Sound | (not specified) | rain on brick, distant traffic, no music, no dialogue |
| Constraints | (not specified) | no on-screen text, no extra limbs, realistic cloth weight |
| Typical result | soft, generic, unusable second half | one clean take, repeatable, gradeable |
Notice that the advanced column is not more poetic. It is more decidable. Every clause resolves to something a render can either do or not do, which is exactly what removes the coin-flip feeling from generation. If you want the long form reasoning behind each of those seven slots, the prompt anatomy breakdown walks through them one at a time with failure examples.
| Category | Most common failure | Clause that fixes it |
|---|---|---|
| Cinematic | over-lit, contrast-free image | name one key source and its direction |
| Ultra realistic | glossy, plastic skin | name the capture device and its flaws |
| Action | speed ramps you did not ask for | state constant real-time speed |
| Sci-fi | screens full of garbled text | no on-screen text, no UI glyphs |
| Horror | jump-cut chaos | one continuous move, hold the frame |
| Nature | impossible animal anatomy | name species, gait, and body weight |
| Anime | 3D render pretending to be 2D | state cel shading and frame rate feel |
| Commercial | logo hallucinations | unbranded packaging, no text |
| Dialogue | lips out of sync | put the exact line in quotes, one speaker |
| Image-to-video | the model redesigns your frame | describe only motion, not appearance |
Cinematic AI video prompts (1–5)
Cinematic does not mean dark and slow. It means someone made decisions about light, lens and blocking before the camera rolled, and the prompt is where you make those decisions. Each of the five below names exactly one key light, one lens, and one move, which is what separates a designed frame from an averaged one. Run these through the cinematic generator and they will hold their contrast instead of flattening out.
Medium shot, 35mm, slow push in at eye level. A woman in her late thirties in a soaked wool overcoat stands still in a narrow brick alley, rain flattening her hair, breath visible. She lifts her chin and holds the look for the length of the shot. Standing puddles reflect a single sodium streetlight positioned behind her, throwing a hard orange rim across her shoulders and leaving her face in deep shadow. A fire door sits propped open in the background. Sound: heavy rain on brick, distant traffic hiss, no music, no dialogue. No on-screen text, no extra limbs, realistic wet cloth weight.
Swap the coat and the streetlight color to reskin this entire shot.
Why it works: The whole frame is built around one light placed behind the subject, which is why it reads as photographed rather than rendered. Backlight gives the model an unambiguous contrast map: bright rim, dark face, reflective ground. The action is deliberately almost nothing, a chin lift, because a still subject lets the rain and the light do the motion work, and rain is one of the few particle systems current models handle beautifully. The wet cloth weight constraint stops the coat from behaving like dry silk, which is the tell that ruins most rain shots.
Wide two-shot, 28mm, static locked-off camera at seated eye level. A man in his sixties and a woman in her thirties sit at opposite ends of a small kitchen table, neither speaking, both looking at the table surface. Late afternoon sun comes through a single window camera-left, hard-edged, cutting a bright rectangle across the tabletop and leaving the rest of the room in falloff. Two cold mugs, a folded newspaper. Sound: refrigerator hum, a clock, one distant car, no music, no dialogue. No on-screen text, no camera movement, no extra limbs.
The locked-off camera is doing the emotional work here. Do not add a push in.
Why it works: Most people reach for movement when a scene feels static, which is backwards. A locked-off wide with two motionless figures creates tension the model cannot accidentally undercut, and it removes every opportunity for the drifting-camera artifacts that appear in slow moves. The hard window light with visible falloff gives depth without any color grading, and specifying that the rest of the room falls off dark prevents the flat overall exposure that video models default to when no light source is named.
Close-up, 85mm long lens with compressed background, static frame with the tiniest handheld drift. A street musician in his fifties, deeply lined face, grey stubble, worn corduroy jacket, exhales and looks slightly off camera. Low golden-hour sun hits him from camera-right at a shallow angle, catching the edge of his cheek and turning the out-of-focus traffic behind him into soft round highlights. Sound: street ambience, a bus passing, no music, no dialogue. Shallow depth of field, background fully defocused. No on-screen text, no face morphing, no zoom.
Change the subject's age and jacket material; keep the 85mm and the golden hour angle.
Why it works: Long lens plus low sun is the most reliable beauty combination available to a text prompt, because the compression flattens facial features in a flattering way and the shallow depth hides background detail the model might otherwise get wrong. Naming the bokeh explicitly as soft round highlights gives the render something concrete to build in the defocused area instead of mush. The no face morphing constraint matters on close-ups specifically, where identity drift is most visible.
Wide establishing shot, 24mm, slow crane up from ground level to twelve feet. A lone figure in a red jacket walks left to right along a black volcanic beach, small in frame, never facing camera. Overcast top light, no visible sun, flat grey sky meeting dark grey water at a high horizon line. Foam retreats across wet black sand. Sound: heavy surf, wind across the microphone, one gull, no music, no dialogue. Camera rises smoothly at constant speed. No on-screen text, no lens flare, no slow motion, realistic water physics.
Red on grey is the only color decision in this shot, which is why it lands.
Why it works: This is a color-restraint shot: one saturated element in an otherwise monochrome frame. Telling the model the sky is flat grey and the sun is not visible removes the golden glow it otherwise applies to every landscape. A crane up at constant speed is easier for the model to sustain than a crane with acceleration, and constant speed also prevents the unrequested speed ramping that plagues moving-camera landscape shots. The high horizon gives the composition somewhere to go as the camera rises.
Medium tracking shot, 35mm, camera walks backwards ahead of the subject at a steady pace. A hotel housekeeper pushes a linen cart down a long carpeted corridor, looking forward, not at camera. Overhead practical sconces every few metres pass over her face in a rhythmic bright-dark-bright pattern. Warm tungsten light, patterned carpet, closed numbered doors receding. Sound: cart wheels on carpet, a distant elevator chime, room ambience, no music, no dialogue. One continuous take, constant walking speed. No on-screen text, no door numbers legible, no extra limbs.
The passing sconces are a free rhythm generator. Any corridor works.
Why it works: Repeating practical lights give a tracking shot built-in visual rhythm and, more usefully, they give the model a periodic exposure pattern to anchor to, which stabilises the whole move. Asking for the door numbers to be illegible is a smarter constraint than banning text outright here, because doors clearly need markings and forcing them blank looks stranger than blurring them. Constant walking speed with a backwards-tracking camera is one of the most reliably clean camera instructions available.
Ultra realistic AI video prompts (6–10)
Realism is not high production value. It is the absence of production value, described precisely. The trick in every prompt below is to name the capture device and then name its flaws, because flaws are what our eyes read as unmediated footage. Send these to the ultra realistic generator, where they route to a native-audio engine and the imperfect ambience arrives with the picture.
Vertical handheld smartphone footage, slightly unsteady, minor auto-exposure hunting as the camera moves. A man in a creased t-shirt makes coffee in a small kitchen, back partly to camera, filmed casually by someone standing in the doorway. Overcast daylight from a window blows out slightly behind him. Ordinary clutter on the counter: a chopping board, a folded tea towel, an open cereal box turned away from camera. Sound: kettle, tap running briefly, room tone, mono phone microphone, no music, no dialogue. No on-screen text, no cinematic grading, no shallow depth of field, no stabilisation.
Banning shallow depth of field is what stops this becoming an advert.
Why it works: Phone cameras have deep focus and clumsy exposure, and stating both is what pushes the render out of cinematic mode. Auto-exposure hunting is a specific, nameable artifact rather than a vague call for realism, and models reproduce it convincingly. Mono microphone is doing the same job on the audio track: without it, native audio tends to arrive as a wide, produced stereo bed that immediately signals fiction. The clutter, described as ordinary and turned away from camera, keeps packaging text out of frame without banning objects.
Fixed dashcam view through a windshield, wide angle, mounted behind the rear-view mirror, faint constant engine vibration. Two lanes of traffic moving at normal speed on a wet motorway at dusk. Wipers cross the frame at intervals, briefly smearing then clearing the view. Tail lights ahead throw red streaks on the wet asphalt. Sky is losing light, headlights on, no sun. Sound: road roar, wipers, faint engine, no music, no dialogue. Constant real-time speed, no slow motion, no camera movement other than vehicle motion. No on-screen text, no timestamps, no overlays, realistic vehicle proportions, matte road surface not mirror-like.
Matte road surface is the clause that prevents the liquid-mirror motorway effect.
Why it works: Dashcam is the single easiest realism archetype to trigger because the framing is so distinctive, but it has two classic failures: the road turning into a mirror and the timestamp overlay appearing uninvited. Both are addressed explicitly here. Naming the wipers gives the shot a periodic event that also proves the camera is behind glass, which sells the archetype more than any adjective. Constant real-time speed blocks the dramatic ramping models like to add to vehicle footage.
Fixed ceiling-mounted security camera, high wide angle looking down at roughly thirty degrees, slight barrel distortion, low frame rate feel. An empty hardware shop floor after closing, aisles of unbranded stock, one row of overhead fluorescents left on and the rest dark. A single employee walks into frame from the lower right, crosses diagonally, and exits upper left without looking up. Sound: fluorescent hum, distant HVAC, footsteps, no music, no dialogue. No on-screen text, no timestamp, no date stamp, no camera movement, unbranded packaging, realistic human gait.
Low frame rate feel is what separates surveillance from a wide-angle film shot.
Why it works: Surveillance realism is mostly about the wrong angle and the wrong frame rate, not grain. Describing the mount height and the downward angle produces a perspective no cinematographer would choose, which is exactly the point. Leaving one bank of lights on and the rest dark is a small, very human detail that reads as after-hours far better than the word night. Unbranded packaging keeps the model from inventing logos across an aisle of products, which is where this kind of shot normally falls apart.
Chest-mounted bodycam, wide angle, heavy natural bob with each step, subject's own forearms occasionally entering the lower frame. Ascending a concrete emergency stairwell lit by cold fluorescent strips on every landing, painted handrail, scuffed steps. Speed is a brisk walk, not a run. Slight lens flare from each strip light as it passes overhead. Sound: breathing, footsteps echoing in concrete, fabric rustle against the microphone, no music, no dialogue. No on-screen text, no timestamps, no third-person shots, no extra limbs, realistic arm movement.
Forearms entering frame is the detail that makes point-of-view believable.
Why it works: Point-of-view footage fails when the model quietly cuts to a third-person angle, which is why banning third-person shots is a necessary clause rather than a redundant one. Letting the subject's own forearms enter the lower frame gives a physical anchor that a floating camera cannot fake. Specifying a brisk walk rather than a run keeps the bob amplitude in a range the model renders cleanly, and fabric rustle on the microphone is the audio equivalent of the forearm detail.
Handheld consumer camcorder footage, 4:3 feel inside the frame, soft video look, mild interlacing artifacts, auto white balance drifting slightly. A small garden party in the middle of a bright afternoon: six adults around a folding table, someone laughing off-frame, the camera operator panning too quickly to follow a child running past and then correcting. Harsh overhead midday sun, hard shadows, blown highlights on white plates. Sound: overlapping voices, birds, wind buffeting the built-in microphone, no music. No on-screen text, no date stamp, no cinematic grading, no slow motion, natural unposed movement.
The overshoot-and-correct pan is a human error no cinematic prompt would ever ask for.
Why it works: Realism lives in operator mistakes. Panning too fast and then correcting is a two-beat behaviour the model can execute and that no professional camera operator would produce, so it lands instantly as amateur footage. Midday sun with blown highlights is another anti-cinematic choice: it is the light everyone avoids, which is why it reads as unplanned. Wind buffeting the built-in microphone completes the illusion, since built-in mics are exactly where wind noise happens.
Action AI video prompts (11–15)
Action prompts fail for one reason more than any other: too much action. A clip has room for one physical beat, and asking for a chase, a jump and a landing gets you a smear of all three. Each prompt below isolates a single beat and then spends its remaining words on weight, friction and impact, which is where believable action actually comes from.
Tracking shot, 35mm, camera moves parallel to the subject at matching speed, slightly behind. A courier in a fitted jacket and backpack sprints across a flat gravel rooftop in a straight line, arms driving, feet visibly displacing loose gravel with each step. Late afternoon sun low from camera-left casting a long running shadow across the roof surface. Air conditioning units in the background. Sound: gravel underfoot, hard breathing, wind, city hum below, no music. Constant real-time speed throughout. No on-screen text, no slow motion, no jump, no extra limbs, realistic body weight and momentum.
Explicitly banning the jump is the difference between one clean sprint and a mangled leap.
Why it works: Models associate rooftop sprints with rooftop jumps and will supply the jump uninvited, ruining the clip halfway through. Naming and banning the second beat is more effective than trying to describe a longer sequence. Gravel is a deliberate surface choice because visible displacement under each footfall is strong physical evidence of weight, and the long shadow gives a second read on speed. Constant real-time speed blocks the speed ramp that action footage attracts.
Close-up, 50mm, static camera slightly below shoulder height, framed on the bag and the striking arm. A boxer's gloved fist lands three times on a worn leather heavy bag, the bag visibly deforming on contact and swinging back between strikes, chain taking up the slack. Dusty gym, single hard overhead light making the bag's texture read. Sweat visible on the forearm. Sound: three heavy leather impacts, chain rattle, breath out on each strike, distant skipping rope, no music. No on-screen text, no slow motion, no full-body shots, realistic impact physics and bag mass.
Naming the number of strikes stops the model from inventing a montage.
Why it works: Impact reads through deformation and recovery, not speed, so the prompt describes the bag compressing and the chain taking up slack rather than asking for a powerful punch. Counting the strikes gives the clip an exact structure that fits inside the duration, which is the same reason it does not drift into a training montage. Keeping the framing tight on the bag and the arm avoids full-body renders where limb count and joint articulation are far more likely to break.
Medium wide, 35mm, static camera at hip height on the pavement side of a parked sedan. A woman in a dark suit steps out of the driver's door, pushes it closed with one hand hard enough that the car body visibly settles on its suspension, and walks out of frame camera-right. Overcast daylight, wet tarmac, no sun. Reflections on the car panel move as the door swings. Sound: door latch, suspension creak, heels on wet tarmac, light traffic, no music, no dialogue. No on-screen text, no visible number plate characters, no slow motion, realistic vehicle weight and door mass.
Suspension settling on the slam is a two-frame detail that convinces instantly.
Why it works: Vehicle shots are where wrong physics is most obvious to a viewer, because everyone has closed a car door. Asking for the body to settle on its suspension gives the render a specific mechanical consequence to produce. Keeping the number plate characters out of frame avoids the garbled alphanumeric mess models generate on plates, and the low hip-height camera makes the car feel heavy without needing a single adjective about weight.
Low angle, 24mm, camera near ground level, static, subject landing toward camera. A skateboarder lands a simple flat-ground trick on rough concrete, board slapping down first, then body weight settling through bent knees, wheels finding grip and rolling on. Bright overcast light, no hard shadows. Chipped concrete, a painted line, weeds at the edge. Sound: board slap on concrete, urethane wheels rolling, a shoe scuff, no music. Real-time speed. No on-screen text, no slow motion, no aerial rotation, no extra limbs, realistic board and body weight.
Board first, then body weight, then roll-on. Three ordered micro-beats inside one action.
Why it works: Sequencing micro-beats inside a single action gives the model a timeline without asking it to render two separate events. Low camera near the ground puts the impact at the centre of the frame and hides the torso, which reduces the surface area for anatomy errors. Naming the specific sound of urethane on concrete gets a much more accurate audio track than the word skateboarding would, and banning aerial rotation keeps the trick inside what the clip length can carry.
Side tracking shot, 85mm long lens, camera moves parallel to the animal at matching speed. A single chestnut horse gallops through ankle-deep water on a wide tidal flat, hooves throwing distinct sheets of spray backwards on each stride, mane and tail moving with the run rather than floating. Low sun behind the camera, warm, backlighting the spray so each sheet catches light. Sound: hooves in water, heavy breathing, wind, no music. Real-time speed, constant gait. No on-screen text, no slow motion, no rider, realistic equine anatomy, four legs, correct galloping gait.
Stating the correct gait explicitly is worth more than any other clause in an animal prompt.
Why it works: Animal locomotion is a known weak point, and the specific fix is naming the gait rather than the action. A gallop has a defined four-beat footfall pattern that models reproduce far better when the word appears. Asking for spray thrown backwards gives directionality to a particle effect that otherwise sprays symmetrically, which is a subtle but common tell. Mane and tail moving with the run rather than floating targets the exact cloth-and-hair physics failure these shots produce.
Sci-fi AI video prompts (16–20)
Science fiction is the category most damaged by unrequested on-screen text. Every prompt below therefore treats screens, panels and displays as light sources rather than information surfaces, which is both more filmable and far more reliable. Beyond that, the rule is to describe unfamiliar things using familiar materials, because a model renders brushed aluminium far better than it renders the word futuristic.
Wide shot, 35mm, extremely slow push in, static horizon. A woman in a worn utility jumpsuit stands with her back to camera at a tall reinforced window inside a small habitat module, hands at her sides, not moving. Beyond the glass, a pale grey planet fills two thirds of the view, its terminator line crossing slowly. Interior is dim, lit only by cold blue light spilling in from the window and one small amber indicator lamp near the frame. Brushed metal walls, exposed conduit, a folded blanket on a bunk. Sound: low ventilation hum, a periodic soft mechanical tick, no music, no dialogue. No on-screen text, no readable displays, no lens flare, no extra limbs.
Two light sources, one cold and one warm, is the entire lighting design.
Why it works: Backs to camera work beautifully in science fiction because they remove face rendering from the equation and put all the model's effort into the environment. The lighting is entirely motivated: cold from the window, one warm accent, nothing else, which is why the interior holds contrast instead of turning into evenly lit grey. Describing walls as brushed metal with exposed conduit gives the render real materials, and banning readable displays keeps the panels as glow rather than gibberish.
Medium wide, 35mm, static camera at ground level looking slightly up. A four-rotor cargo drone the size of a small car descends slowly into a wet industrial yard, downwash flattening standing water into circular ripple patterns and pushing loose debris outward. Matte grey composite body, visible rivets, one steady red navigation light. Night, lit by two overhead floodlights on poles. Sound: rotor wash building, water spray, metal creak on touchdown, no music, no dialogue. Constant descent speed. No on-screen text, no logos, no readable markings, no lens flare, realistic downwash and water physics.
Downwash on standing water is what proves the aircraft has mass.
Why it works: Flying vehicles look weightless unless they visibly affect their surroundings, so the prompt spends most of its words on the interaction rather than the object. Circular ripple patterns and outward debris are specific, renderable consequences. Matte grey composite with visible rivets keeps the model from producing a glossy concept-art object, and banning readable markings prevents the hull text that otherwise appears on every vehicle panel.
Macro shot, static, extremely shallow depth of field, subject centred. A translucent bioengineered specimen the size of a plum sits suspended in clear fluid inside a sealed cylindrical container, its internal structures faintly pulsing at a slow steady rhythm. Cold white light from directly below the container passes up through the fluid, making the internal structures glow from within. Surrounding lab is dark and out of focus. Sound: low electrical hum, fluid movement, a slow soft pulse, no music, no dialogue. No on-screen text, no labels, no readable instruments, no camera movement, no rapid motion.
Underlighting through fluid is a lighting setup, not an effect, and models render it accurately.
Why it works: Macro plus a single unusual light direction produces images that feel authored rather than generated. Light passing up through fluid is physically simple and visually striking, and the model has plenty of reference for it. Naming the pulse as slow and steady prevents the frantic flickering that ambiguous words like pulsing energy produce. Keeping the lab dark and defocused removes an entire environment the render would otherwise have to invent, and with it a dozen chances to hallucinate labels.
Wide shot, 28mm, static camera, high angle from a mezzanine. Around thirty commuters wait on a clean underground transit platform, most standing still, a few shifting weight, none looking at camera. Continuous soft light strips run along the ceiling and the platform edge, cool white, evenly distributed with no hard shadows. Pale concrete, a curved tunnel mouth, glass barrier doors. Sound: ventilation, low crowd murmur, a soft approaching rumble, no announcements, no music, no dialogue. No on-screen text, no signage, no advertising panels, no readable displays, no extra limbs.
Banning announcements as well as text keeps both channels clean.
Why it works: Crowd scenes are risky, so the prompt reduces motion to almost nothing and shoots from a high angle where faces are small and limb errors are less visible. Continuous light strips are a strong near-future cue that requires no invented technology. The audio constraint is unusual here in banning spoken announcements specifically, because native audio models love station announcements and will produce garbled speech that draws attention to itself far more than any visual flaw.
Tight close-up, 85mm, static, framed on a helmet visor filling most of the frame. A gloved hand and a section of white composite structure are reflected across the curved gold-tinted visor of an environment suit as the wearer works on something just below frame. The face behind the visor is only faintly visible. Hard unfiltered key light from camera-right, no atmospheric diffusion, black background with no stars in frame. Sound: amplified breathing inside the helmet, muffled fabric and mechanical contact, radio static, no music. No on-screen text, no heads-up display, no readable indicators, no extra limbs, no face morphing.
Curved reflective surfaces give the model a legitimate reason to show detail it never has to render properly.
Why it works: Reflections on a curved visor are forgiving: they can be impressionistic and still look correct, so the shot gets visual complexity for free. Hard light with no diffusion and a black background is physically right for the setting and also removes the atmospheric haze the model applies by default. Banning the heads-up display is the sci-fi version of banning on-screen text, and it is the single most common thing that ruins helmet shots.
Horror AI video prompts (21–25)
Horror is the category where restraint pays the largest dividend, because the scariest frame is usually the one where very little happens and the camera refuses to cut away. Prompts that ask for a monster get a rubber monster. Prompts that ask for one wrong detail in an ordinary room get something far worse. Every example below holds a single frame and lets duration do the work.
Medium wide, 35mm, static locked-off camera in the middle of a domestic hallway at night. A bedroom door at the end of the hall stands open about ten centimetres, the gap completely black. Nothing enters or exits. The only light is a weak amber landing lamp behind the camera, throwing the camera's own faint shadow along the carpet toward the door. Floral wallpaper, a radiator, a coat on a hook. Sound: house settling, a fridge cycling on in another room, no music, no dialogue, no footsteps. No on-screen text, no camera movement, no figures, no jump scare, nothing enters the frame.
Explicitly telling the model that nothing happens is what makes it work.
Why it works: Video models are trained on footage where something occurs, so they will manufacture an event unless told not to. Naming the absence, no figures and nothing enters the frame, produces the sustained dread that a described monster never achieves. The camera's own shadow reaching toward the door is a compositional detail that implies a presence behind the lens without rendering one. Static framing eliminates every artifact that camera motion introduces.
Close-up, 50mm, static, subject slightly off centre. An elderly woman sits close to a small open fire, the flame the only light source, positioned low and camera-left so it lights her face from beneath and leaves the top of her head in darkness. She stares slightly past the camera and does not blink for the length of the shot. Firelight flickers unevenly across her features. Sound: fire crackle, wind outside, no music, no dialogue. No on-screen text, no camera movement, no other figures, no face morphing, no rapid movement, realistic skin texture.
Uplighting from a low practical is the oldest horror lighting trick and models execute it perfectly.
Why it works: Light from below inverts the shadow pattern our brains expect on a human face, which produces unease without any monster design. Firelight is the ideal source because its natural flicker adds motion to an otherwise still frame, meaning the clip never feels frozen even though nobody moves. Not blinking is a single, precise behavioural instruction that costs nothing to render and does more than any expression adjective. Realistic skin texture guards against the waxy look that close-up faces drift toward in low light.
Medium shot, 50mm, static camera behind and slightly to the side of the subject. A man in his forties brushes his teeth at a bathroom mirror, ordinary morning behaviour, harsh overhead fluorescent light. His reflection is visible in the mirror and moves in perfect sync with him. The bathroom door behind his reflection is closed. Nothing else is in the room. Sound: tap dripping, brushing, extractor fan, no music, no dialogue. Real-time speed, one continuous shot. No on-screen text, no camera movement, no additional figures, no jump scare, correct mirror reflection geometry.
The horror here is entirely in the setup and the duration. Do not add a figure.
Why it works: This prompt deliberately describes a completely normal scene with one loaded piece of framing, the closed door visible in the reflection, and lets the viewer supply the dread. Asking for correct mirror geometry is a genuinely useful technical constraint, because mirrors are a known failure case and an incorrect reflection reads as a cheap glitch rather than a scare. Harsh overhead fluorescent is the ugliest available light, which is exactly right for the bathroom.
Handheld point-of-view, wide angle, natural walking bob, the only light a single handheld torch beam pointing where the camera looks. Dense conifer woodland at night, wet ground, low fog sitting between trunks at knee height. The beam picks out bark, ferns, and then dark space between trees where the light falls off completely. The walk is slow and steady, never running. Sound: footsteps on wet needles, breathing, distant branch movement, no music, no dialogue, no screaming. No on-screen text, no figures in the beam, no jump scare, realistic light falloff.
Realistic light falloff is a technical clause with a huge atmospheric payoff.
Why it works: A torch in the woods is a lighting instruction disguised as a location. Requiring realistic falloff forces the render to let the beam die into blackness rather than gently illuminating the whole forest, and that blackness is the source of the tension. Fog at knee height gives the beam something to scatter through, which makes the light shaft visible and physically legible. Banning figures in the beam keeps the anticipation intact for the length of the shot.
Wide shot, 28mm, static camera at the shallow end of an indoor swimming pool at night. The pool is full and completely still, surface like glass, no swimmers. Underwater lights are on, throwing rippling caustic patterns on the tiled walls and the high ceiling even though the water is not moving. Everything else is dark. Wet tiles, a stack of floats, a lifeguard chair. Sound: water filtration hum, a single distant drip echoing in the tiled space, no music, no dialogue, no splashing. No on-screen text, no camera movement, no figures, no ripples on the surface.
Still water plus moving caustics is a contradiction the render resolves beautifully.
Why it works: Asking for a glass-still surface while keeping caustic patterns on the walls creates a subtly impossible image that nobody consciously flags but everybody feels. Underwater lighting in a dark room is also a strong contrast structure the model handles well. The audio design carries most of the dread: an echoing drip in a large tiled space is one of the most specific and reproducible sound cues you can request, and banning splashing keeps the water empty.
Fantasy AI video prompts (26–30)
Fantasy prompts collapse when they ask for magic. Magic is not a visual, it is a category, and the model averages every glowing effect it has ever seen. The prompts below always specify what the magic physically does to light, air, dust or water, because that is renderable. Describe the consequence, never the concept.
Medium close-up, 50mm, static camera, framed on an anvil and a pair of forearms. A blacksmith brings a hammer down once on a glowing orange billet, sending a burst of small sparks arcing away and dying within half a metre. The billet dims slightly at the strike point. The forge behind is the only light source, deep orange from camera-right, leaving the workshop in near darkness. Soot, hanging tongs, a quenching barrel. Sound: one heavy metal impact with a long ring-out, fire roar, sparks ticking on stone, no music, no dialogue. No on-screen text, no full-body shot, no slow motion, realistic spark physics and short spark lifetime.
Short spark lifetime stops the render turning into a firework display.
Why it works: Sparks are the effect models most reliably overdo, so bounding them in both distance and duration is the key clause. One strike with a ring-out gives the clip a clean acoustic and visual shape that fits inside any duration. Lighting the entire frame from the forge alone means the render has one unambiguous source, and the dimming at the strike point is a small physical consequence that makes the metal feel hot rather than merely orange.
Wide shot, 35mm, extremely slow push in. A vast abandoned library with collapsed upper galleries and books scattered across a stone floor. One shaft of hard daylight enters through a hole in the roof and lands on a single reading table, thick with visible dust particles drifting slowly through the beam. Everything outside the beam is deep shadow. No people. Sound: a distant drip, faint wind moving through the roof, settling debris, no music, no dialogue. No on-screen text, no readable book titles, no figures, realistic dust motion, no lens flare.
Dust in a light shaft is the cheapest atmosphere in cinema, and models render it accurately.
Why it works: A single hard beam in a dark space gives the frame both subject and structure, and airborne dust turns the beam into a visible volume rather than a bright patch. The slow push in adds parallax between foreground rubble and the beam, which sells the scale of the room. Banning readable titles keeps the model from covering every spine with hallucinated typography, which is the failure that instantly cheapens library shots.
Wide shot, 85mm long lens compressing the distance, static camera. A single cloaked rider on a dark horse waits motionless at the edge of a pine treeline, small in the lower third of the frame, facing away from camera into the trees. Cold blue pre-dawn light, heavy ground mist reaching the horse's knees, no sun visible. The cloak moves slightly in a light wind. Sound: wind through pines, the horse shifting weight, one distant bird, no music, no dialogue. No on-screen text, no face visible, no weapons drawn, realistic equine anatomy, realistic cloth weight.
Long lens plus ground mist compresses depth into flat layers, which is why it reads as painted.
Why it works: Compressing a landscape with a long lens stacks the mist layers into distinct planes, an effect that looks deliberate and expensive. Keeping the rider small, motionless and turned away removes faces, hands and reins from the render, which are exactly the elements that break. Cold pre-dawn with no visible sun overrides the warm golden default that landscape prompts attract, and the slight cloak movement keeps the frame alive without adding action.
Close-up, 50mm, static camera looking down at a shallow stone basin. Clear water in the basin rises slowly into a smooth dome about ten centimetres high and holds that shape, surface tension intact, no droplets escaping, then settles back flat. Cold overcast daylight from above, no direct sun. Wet moss on the stone rim. Sound: a low sustained tone, water movement, dripping at the edges, no music, no dialogue. Slow steady motion throughout. No on-screen text, no hands, no sparkles, no glowing particles, realistic water surface tension.
Banning sparkles is what keeps this magical instead of cartoonish.
Why it works: This is the clearest demonstration of describing a consequence rather than a concept: there is no mention of magic anywhere, only water doing something water cannot do. The model renders it convincingly because every individual element, a dome of water with intact surface tension, is physically describable. Sparkles and glowing particles are the default visual shorthand for magic and banning them forces the render to keep its realism, which is what makes the impossible part land.
Low wide angle, 24mm, static camera at ground level in tall grass. A vast quadruped creature the size of a house walks slowly past in the middle distance from left to right, only its legs and lower body in frame, each footfall pressing the grass flat and sending a faint tremor visible in nearby grass stems. Overcast diffuse daylight. Grass in the immediate foreground is sharp and slightly out of focus at the edges. Sound: deep low-frequency footfalls, grass movement, wind, distant birds, no music, no roaring. No on-screen text, no full creature visible, no camera movement, realistic weight and slow gait.
Keeping the creature partly out of frame is both better filmmaking and a safer render.
Why it works: Cropping the creature at the top of frame does two things at once: it communicates scale far better than showing the whole animal, and it hides the head, which is where anatomy errors would be most obvious. The tremor in nearby grass gives mass a visible consequence. Banning roaring is deliberate, since native audio will otherwise supply a stock monster roar that undercuts the restraint of the visual.
Nature and wildlife AI video prompts (31–35)
Wildlife prompts live or die on two words: the species and the gait. Naming both gives the model a specific locomotion pattern to reproduce; leaving either vague produces the boneless animal motion that gives AI footage away. Documentary framing helps too, because a long lens from a distance is both the honest way to shoot animals and the safest way to render them.
Medium shot, 300mm long lens, static camera at ground level, heavily compressed background. A single red fox trots across fresh unbroken snow from right to left in a straight line, four legs in a clear trotting gait, tail held low and level, leaving a distinct line of prints behind it. Flat overcast winter light, no shadows, no sun. Bare birch trunks far behind, fully defocused. Sound: paws compressing snow, faint wind, no music, no vocalisation. Real-time speed. No on-screen text, no slow motion, no humans, realistic vulpine anatomy and trotting gait, four legs only.
The trail of prints behind the animal is proof of contact with the ground.
Why it works: Footprints left in snow are the single most convincing detail available in a wildlife shot, because they demonstrate that the animal's feet are actually meeting the surface rather than sliding above it. Naming the gait as a trot and the tail position as low and level gives the model two precise skeletal constraints. Flat overcast light removes shadow rendering entirely, which is one fewer thing to get wrong, and the long lens keeps the background a soft wash.
Wide underwater shot, slow upward tilt, natural buoyancy drift in the camera. Tall kelp fronds rise from the seabed toward a bright surface, swaying together in a slow surge rhythm rather than randomly. Shafts of sunlight cut down through the water column between the fronds, scattering in suspended particles. Small silver fish move in a loose group in the mid distance. Sound: muffled underwater ambience, distant surge, no music, no bubbles close to camera. No on-screen text, no divers, no camera lights, realistic underwater light attenuation and colour loss with depth.
Surge rhythm, not random motion, is what makes underwater plants look real.
Why it works: Kelp does not wave individually, it moves as a mass on the swell, and specifying that produces coherent motion instead of noodly independent fronds. Requesting realistic light attenuation gets the correct blue-green colour shift with depth, which is otherwise ignored and leaves the shot looking like a fish tank. Light shafts scattering in suspended particles give the water column visible volume, exactly as dust does in air.
Very wide static shot, 24mm, locked-off camera on flat open grassland with a low horizon line and a huge sky. A single storm cell crosses from left to right in the far distance, visible rain shaft hanging beneath it, the base flat and dark, the top bright where sun still reaches. Sunlight falls on the foreground grass while the distance is in shadow. Grass moves in a steady wind toward camera. Sound: wind across grass, one distant low thunder roll, no music. Real-time speed, no time lapse. No on-screen text, no lightning strikes in frame, no camera movement, realistic cloud motion speed.
Banning time lapse keeps the clouds moving at a believable rate.
Why it works: Sky prompts almost always come back accelerated, because most training footage of dramatic clouds is time lapse. Explicitly requesting real-time speed and realistic cloud motion produces a much rarer and more cinematic result: a sky that is almost, but not quite, still. Splitting the light so the foreground is sunlit and the distance is in shadow gives the frame depth, and keeping lightning out of frame avoids the flicker artifacts strikes generate.
Macro shot, 100mm macro lens, static camera, extremely shallow depth of field. A hummingbird hovers at a red tubular flower, wings a genuine blur at real-time speed, body held remarkably stable, beak entering the flower and withdrawing once. Soft directional morning light from camera-left, background a smooth green wash completely out of focus. Sound: rapid wingbeat hum, faint garden ambience, no music. Real-time speed. No on-screen text, no slow motion, no frozen wings, realistic avian anatomy, stable hovering body with only the wings blurred.
Asking for blurred wings and a stable body prevents the whole bird from smearing.
Why it works: The signature look of a hovering hummingbird is a sharp body with unresolvable wings, and stating that split explicitly stops the model from either freezing the wings or blurring the entire animal. Banning slow motion is essential here because almost all reference footage of hummingbirds is high speed. Macro plus a smooth defocused background removes the environment from the render, concentrating quality where the eye goes.
Overhead close shot, 50mm, static camera looking straight down into a shallow rock pool. Clear water about ten centimetres deep over barnacled rock, one anemone slowly closing, small shrimp moving in short bursts, the surface disturbed occasionally by a light breeze producing brief ripple patterns and refraction across the rock below. Bright overcast light, no direct sun, no reflection of the sky. Sound: distant surf, water lapping, no music. Real-time speed. No on-screen text, no hands, no camera movement, realistic water refraction and slow marine animal movement.
Refraction through moving water is a physics test models now pass, so use it.
Why it works: Looking straight down through shallow rippling water is a real technical showcase: the render has to distort the rock beneath in step with the surface. Modern models do this well, and it produces a shot that is hard to dismiss as generated. Specifying no reflection of the sky is what keeps the pool transparent instead of turning into a mirror, and asking for short-burst shrimp movement gives the small animals a believable motion signature.
Anime and stylised animation prompts (36–40)
The commonest anime failure is a 3D render wearing a 2D costume: correct colours, correct hair, but smooth interpolated motion that no animation studio would produce. The fix is to describe the production technique, not just the look. Say cel shading, say flat colour fills, say limited animation on twos, and the render changes character completely.
Hand-drawn 2D anime style, cel shading with flat colour fills and clean dark line art, limited animation on twos, no 3D rendering. Medium shot, static frame with a very slow horizontal pan. A high school student in a plain uniform stands at a rooftop railing looking out over a city, hair and tie moving in a light wind. Deep orange sunset sky with painted cloud bands, warm rim light on the character's shoulders, long shadow across the rooftop concrete. Sound: wind, distant traffic, a school bell far away, no music, no dialogue. No on-screen text, no subtitles, no watermark, consistent character design throughout.
Naming animation on twos is what removes the uncanny smoothness.
Why it works: Traditional television animation holds each drawing for two frames, which produces a distinctive stutter our eyes read instantly as hand drawn. Requesting it directly is far more effective than any list of visual adjectives. Flat colour fills and clean line art push the render away from painterly 3D shading, and a painted sky is a specific reference to how anime backgrounds are actually produced. Banning subtitles matters because anime training data is saturated with them.
2D anime style, cel shaded, detailed painted background with simplified character art. Wide static shot from across a wet street at night, looking at the bright interior of a small convenience store through its glass front. A single figure stands inside at the counter, seen in silhouette against the fluorescent interior. Rain falls steadily, streetlight reflections stretch across the wet road, neon colour spills onto the puddles. Sound: rain, a distant train, the store's automatic door chime, no music, no dialogue. No on-screen text, no signage lettering, no subtitles, no watermark, consistent line weight.
Silhouetting the character removes face consistency from the equation entirely.
Why it works: Bright interiors against dark wet streets are a signature of the style and a strong contrast structure for any render. Putting the only character in silhouette is a compositional decision that also solves the hardest technical problem in stylised video, which is keeping a drawn face on-model across frames. Asking for detailed backgrounds with simplified characters mirrors real production practice, and consistent line weight is a surprisingly effective clause against drifting art quality.
2D anime action style, cel shaded, high contrast, dynamic linework. Medium shot, camera whip-pans to follow a single motion. A cloaked figure lands hard in a stone courtyard from above, dust bursting outward radially on impact, radial speed lines emphasising the drop, one held impact frame at the moment of landing before motion resumes. Overcast grey sky, cool colour palette, hard shadows. Sound: wind rush, a single heavy impact, stone debris, no music, no dialogue, no voice. Real-time speed. No on-screen text, no subtitles, no watermark, no extra limbs.
The held impact frame is a native animation technique and it survives the render.
Why it works: Animation punctuates action with a briefly held frame at the moment of impact, and asking for it produces a rhythm no realistic prompt would ever generate. Radial speed lines and radial dust are drawn effects rather than simulated ones, so naming them keeps the shot inside the stylised idiom. The cloak, hood and downward angle keep the face mostly hidden, which is again the safest way to render a stylised character in motion.
Hand-drawn 2D animation, soft cel shading, warm muted palette, painted background, gentle limited animation. Static medium shot of a small kitchen in the morning. An older woman pours tea from a heavy pot into two cups on a wooden table, steam rising, a cat asleep on a chair in the background. Warm morning light through a lace curtain camera-left, dappled on the table surface. Sound: pouring water, a ticking clock, birds outside, no music, no dialogue. Real-time speed. No on-screen text, no subtitles, no watermark, consistent character design, natural hand movement.
One domestic action, rendered patiently, is the entire genre.
Why it works: Quiet animation is defined by its willingness to spend a whole shot on pouring tea, and prompts that respect that get much better results than prompts asking for wonder. Steam and dappled curtain light are two soft, forgiving effects that add motion without stressing the render. Keeping the camera static means every frame of movement belongs to the character, which is where limited animation looks its best.
Stylised 3D animation, soft global illumination, chunky simplified geometry, matte clay-like surfaces, no photorealism, no subsurface skin detail. Medium shot, gentle arc of the camera around the subject. A small round robot with two large eye lenses wobbles as it takes three careful steps across a smooth pastel-coloured floor, catching its balance on the third step. Soft even studio lighting with a coloured bounce from camera-right. Sound: small servo movements, three soft footfalls, a gentle chime on the wobble, no music, no dialogue. No on-screen text, no logos, no realistic textures, consistent proportions throughout.
Matte clay surfaces are what keep stylised 3D from drifting toward photorealism.
Why it works: Stylised 3D collapses into generic CGI unless the surface treatment is specified, and clay-like matte materials with soft global illumination is a precise, well-represented look. Counting the steps and naming the wobble on the third gives the animation a beat structure. Consistent proportions is the stylised equivalent of no face morphing, and it does real work over the length of a clip.
Commercial and product AI video prompts (41–45)
Product prompts have one unique enemy: typography. Any packaging, label or device screen invites the model to invent text, and invented text is always wrong. Everything below keeps surfaces unbranded and lets light, material and motion carry the sell. The result is footage you can drop your real logo onto in post, which is what you wanted anyway.
Studio product shot, 100mm, static camera at label height, extremely shallow depth of field. A tall matte black glass bottle with no label, no text and no logo rotates slowly on a turntable against a seamless dark grey background. A single large soft key light from camera-right creates one long vertical highlight down the bottle's edge, and a narrow strip light from behind-left adds a second thin rim. Slow constant rotation. Sound: a low room tone, no music, no dialogue. No on-screen text, no labels, no branding, no reflections of a studio or crew, no dust, realistic glass refraction.
Two lights, one soft and one hard, is the entire language of product photography.
Why it works: Glass is defined by its highlights, so a product prompt should specify the lights rather than the object. A broad soft source plus a narrow rim gives the render exactly the two shapes it needs to draw. Banning crew reflections is a genuinely useful clause, because reflective products regularly come back with a ghostly figure in the surface. Constant rotation is easier to sustain than any camera move around a static object.
Macro shot, 100mm macro, camera moves slowly left to right across the surface, extremely shallow depth of field. A heavy woven wool fabric in deep charcoal fills the frame, individual fibres catching the light. A hard light rakes across the surface from camera-left at a very shallow angle, exaggerating the weave texture and casting tiny shadows in every valley of the cloth. Sound: faint fabric movement, room tone, no music, no dialogue. Slow constant camera speed. No on-screen text, no labels, no branding, no hands, no wrinkles appearing or disappearing, realistic fibre detail.
Raking light at a shallow angle is a texture microscope, and it costs nothing.
Why it works: Texture only exists on camera when light comes across it rather than at it, and stating the angle explicitly is what produces the tiny shadow in every weave valley. This is the highest-value technique in material advertising and it is almost never specified in prompts. A slow lateral move gives the shot progression without introducing perspective change, which keeps the macro focus plane stable.
Overhead shot, 50mm, static camera looking straight down at a light wooden table. Two hands lift the lid from a plain matte white box and set it aside, revealing folded tissue paper inside. The movement is unhurried and deliberate, both hands visible and in frame throughout. Soft diffused daylight from camera-left, gentle shadow under the box. Sound: cardboard friction, tissue paper, a soft set-down on wood, no music, no dialogue. Real-time speed. No on-screen text, no logos, no branding, no printed patterns, exactly two hands, five fingers per hand, realistic hand anatomy.
Counting fingers out loud in the constraint line still measurably helps.
Why it works: Hands remain the most scrutinised element in any product video and the most likely to fail, so the prompt keeps both hands fully in frame where errors would be obvious, and then constrains them numerically. Overhead framing removes the face and the room. Unhurried deliberate motion is easier to render cleanly than quick action, and it also happens to be how premium unboxing footage is actually shot.
Close-up, 85mm, static camera at glass height, shallow depth of field. A clear amber liquid pours in a steady unbroken stream into an empty clear glass on a dark surface, the level rising smoothly, a small controlled head of bubbles forming and settling at the surface. Backlit by a soft panel directly behind the glass, making the liquid glow and the edges of the glass read as bright lines. Sound: liquid pouring into glass, pitch rising slightly as the glass fills, no music, no dialogue. Real-time speed. No on-screen text, no branding, no hands, no splashing over the rim, realistic fluid dynamics and surface tension.
The rising pitch as the glass fills is a real acoustic phenomenon and native audio reproduces it.
Why it works: Backlighting a transparent liquid is the standard beverage technique because it turns the product into its own light source. The prompt asks for one continuous unbroken stream, which is far more stable than a splashy pour and avoids the droplet chaos models produce. The audio note is not decorative: filling vessels genuinely change pitch, and asking for it produces a soundtrack that feels recorded rather than assembled.
Extreme close-up, 100mm macro, extremely slow push in, shallow depth of field. A brushed steel wristwatch case and bracelet on a dark slate surface, the crown and the case edge in sharp focus, the dial mostly out of frame and unreadable. A soft light sweeps slowly across the brushed metal, travelling along the grain and making the finish shift from dark to bright. Sound: room tone, a faint mechanical tick, no music, no dialogue. Slow constant motion throughout. No on-screen text, no numerals, no dial markings, no logos, no branding, realistic metal reflection.
Framing the dial out of shot is the cleanest way to avoid hallucinated numerals.
Why it works: Watch faces are a typography minefield, so the composition solves the problem before the constraint line has to. A moving light across brushed metal is the definitive way to show that finish, because the grain only reveals itself as the highlight travels. That travelling highlight also provides all the motion the shot needs, which means the camera barely has to move and the macro focus stays where you put it.
Food and drink AI video prompts (46–50)
Food footage is a physics problem wearing an appetite costume. Steam, melt, sizzle and pour are all continuous deformations, and each has a characteristic speed that the prompt has to name or the render invents. State the timescale, name the material, and keep hands to a minimum.
Close-up, 50mm, static camera at a low angle just above the pan rim, shallow depth of field. A knob of cold butter slides across the surface of a hot cast iron pan, immediately beginning to melt at the edges, foaming lightly, and leaving a widening pool of clear golden fat behind it. Warm overhead kitchen light from camera-right, one bright specular highlight on the melted surface. Sound: an immediate sharp sizzle settling into a steady crackle, extractor fan, no music, no dialogue. Real-time speed. No on-screen text, no hands, no branding, realistic melting rate and fluid behaviour.
Sizzle that starts sharp and settles is the audio signature of hot metal.
Why it works: Melting is a rate, and stating a realistic melting rate keeps the butter from vanishing in half a second or sitting inert. The trail of clear fat behind the sliding butter is a specific, observable detail that gives the render a physical narrative inside a two second window. The audio instruction has a shape rather than a label, which is how you get native audio that matches the picture instead of running underneath it.
Macro shot, 100mm, static camera level with the portafilter spouts, shallow depth of field. Two dark streams of espresso emerge from a bottomless portafilter, starting slow and thin, thickening into a steady golden-brown flow, falling into a small white cup below. Backlit by a warm light source behind and slightly below, making the crema glow at the surface. Steam rises faintly. Sound: a low pump hum, liquid hitting ceramic then hitting liquid as the cup fills, no music, no dialogue. Real-time speed. No on-screen text, no logos, no hands, realistic viscosity and crema formation.
Liquid hitting ceramic then hitting liquid is a two-stage sound almost nobody specifies.
Why it works: Extraction has a natural arc, thin to thick, and describing that arc gives a short clip a beginning and an end. Viscosity is the property that separates espresso from coffee-coloured water, and naming it changes the flow behaviour visibly. The two-stage audio detail is small but it is exactly the kind of specificity that makes a native audio track feel synchronised rather than dubbed.
Close-up, 50mm, static camera, shallow depth of field. Two hands tear a round crusty sourdough loaf in half, the crust cracking audibly and shedding a few crumbs, the open crumb structure revealing an irregular honeycomb of holes with visible steam escaping from the hot interior. Soft window light from camera-left, a dark wooden board beneath. Sound: crust cracking, a soft tear, crumbs falling on wood, no music, no dialogue. Real-time speed. No on-screen text, no branding, exactly two hands, five fingers per hand, realistic steam behaviour and crumb structure.
Irregular honeycomb is a much better instruction than the word artisan.
Why it works: Describing the internal structure of the bread gives the render a concrete pattern rather than a marketing adjective, and irregular is the key word since machine-perfect crumb reads as fake immediately. Steam is one of the most reliable appetite cues available and it also proves the bread is hot, a fact no adjective can establish. Hands are unavoidable here, so they are constrained numerically instead.
Close-up, 85mm, static camera slightly above the subject, shallow depth of field. A single scoop of vanilla ice cream in a plain cone begins to melt at the edges, one drip forming slowly, running down the side of the cone and being absorbed into the wafer. Bright hard afternoon sunlight from camera-right, strong shadow, a slightly overexposed background of blurred green foliage. Sound: distant outdoor ambience, faint birdsong, no music, no dialogue. Real-time speed, slow melting rate. No on-screen text, no hands, no branding, no slow motion, realistic melting and absorption.
Absorption into the wafer is the detail that turns a drip into a story.
Why it works: Following a single drip from formation to absorption gives the clip a complete micro-narrative that fits comfortably in a few seconds. Hard sunlight is both physically correct for melting and visually punchy, and the slightly overexposed background is an honest consequence of exposing for the ice cream. Banning slow motion is necessary because almost all reference footage of drips is high speed.
Macro shot, 100mm, static camera at board level, extremely shallow depth of field. A very sharp knife passes cleanly through a ripe tomato in one continuous stroke, the skin giving way without crushing, seeds and gel visible at the cut face, a small amount of juice released onto the wooden board. Bright soft light from directly above, no hard shadows. Sound: the specific quiet split of skin, the blade meeting the board once, no music, no dialogue. Real-time speed. No on-screen text, no branding, no hands above the wrist, realistic cutting physics, no crushing or deformation of the fruit.
No crushing is the clause that gets you a sharp knife instead of a blunt one.
Why it works: The difference between appetising and unpleasant food footage is almost always whether the ingredient is being cut or squashed, so the prompt states it directly. Keeping hands out of frame above the wrist removes most of the anatomy risk while retaining the human element. One continuous stroke, one board contact, one audio event: the whole shot is a single clean beat.
Social and short-form AI video prompts (51–55)
Short-form is not landscape footage cropped narrow. The subject has to live in the middle third of a 9:16 frame, the first half second has to resolve into something recognisable, and there is no time for a camera move that only pays off at second four. Every prompt below states the vertical ratio inside the prompt itself, because models frame differently when they know the shape in advance. If you publish vertically often, vertical video that holds attention covers the composition rules these prompts are built on.
Vertical 9:16 framing, close-up, 50mm, static camera at counter height. A pair of hands pours milk into a dark espresso in a white ceramic cup held centre frame, one continuous pour that finishes inside the shot. Steam rises. Soft north-facing window light from camera-left, gentle falloff, no hard shadows, matte grey concrete counter. Subject occupies the middle third of the vertical frame with clear headroom above and below. Sound: milk hitting liquid, a faint cafe murmur, no music, no dialogue. No on-screen text, no watermarks, no slow motion, realistic liquid physics.
Anything with a start and an end inside two seconds works as a vertical hook. The pour is just the cleanest one.
Why it works: Vertical clips are judged in the first half second, so the action has to be legible immediately and complete before the viewer decides. A pour does both: it reads instantly and it visibly finishes. Stating the 9:16 ratio and the middle-third placement inside the prompt keeps the model from composing for landscape and letting the cup drift toward an edge that will later be covered by interface elements. Real-time speed matters more here than anywhere else, because unrequested slow motion turns a two-second hook into a four-second scroll-past.
Vertical 9:16, medium close-up, 35mm, slow handheld drift. A person's hands, forearms only, sort a small stack of index cards on a wooden desk, laying three of them down one at a time. No face in frame at any point. Warm desk lamp from camera-right plus soft daylight fill from behind camera, shallow depth of field, background bookshelf fully defocused. Sound: paper on wood, a chair creak, a room tone bed, no music, no dialogue. No on-screen text, no legible writing on the cards, no extra limbs, real-time speed.
Hands-only b-roll is the safest AI footage you can cut into a real talking-head video.
Why it works: Faces are where AI footage gets caught, so a hands-only frame removes the single biggest tell while still delivering human presence. Asking explicitly for the card writing to be illegible is better than banning text: cards obviously have writing on them, and a blank card looks wrong, whereas defocused marks look correct. The slow handheld drift is deliberate imperfection, and it is what lets this cut sit beside real camera footage without announcing itself.
Vertical 9:16, close-up, 50mm, camera static while the subject moves. A hand lifts an unbranded matte black cylindrical bottle from a stone surface into centre frame, rotates it ninety degrees, and holds it still. Soft large-source light from camera-left creating one long soft highlight down the side of the bottle, deep neutral grey background falling into shadow. Sound: a single object placement click, quiet room tone, no music, no dialogue. Unbranded packaging, no labels, no logos, no on-screen text, no reflections of a camera crew, realistic weight in the wrist.
Add your real product label in post. Never ask the model to render one.
Why it works: Every hallucinated logo in AI product footage comes from the same mistake: asking the model to render branding. Specify unbranded and add the label afterwards. The single long soft highlight is the whole lighting design, and naming it means the model builds one large source instead of scattering speculars everywhere. Realistic weight in the wrist is a small clause that fixes the floating, weightless handling that makes most AI product shots feel synthetic.
Vertical 9:16, medium shot, 35mm, handheld follow at walking pace. A teenager in a yellow raincoat crosses a wet crosswalk toward camera, steps up onto the kerb, and stops. Late overcast afternoon, flat soft top light, no visible sun, reflections on wet asphalt, blurred traffic behind. Subject centred in the middle third with the head well clear of the top edge. Sound: rain on fabric, tyres on wet road, a distant horn, no music, no dialogue. No on-screen text, no slow motion, realistic cloth and water behaviour.
Cross, step, stop. Three beats, one shot, no cut needed.
Why it works: Short-form rewards a complete micro-action, and cross-step-stop is a full arc in under three seconds. Handheld follow at walking pace matches the speed of the subject, which prevents the sliding-camera artifact that shows up when a move outpaces the action it is tracking. Flat overcast light is the most forgiving lighting for a moving subject because there is no hard shadow to smear as the body turns.
Vertical 9:16, medium shot, 85mm, locked-off camera. A single length of deep red silk hangs from an unseen point at the top of frame and moves in a steady side wind, never fully leaving frame, motion identical at the start and end of the clip. Bright diffuse daylight from behind the fabric making it glow translucent, plain out-of-focus sand-coloured wall behind. Sound: light wind, fabric snap, no music, no dialogue. Continuous constant wind, no gusts, no on-screen text, realistic cloth physics, no slow motion.
Constant wind, not gusts. That single word is what makes it loopable.
Why it works: Loops need motion without narrative, and a constant force gives the model something to sustain evenly across the clip rather than building to a peak. Backlighting translucent fabric is the cheapest way to get genuinely beautiful frames out of a video model, because it turns the material itself into the light source. Saying no gusts removes the acceleration that would make the first and last frame irreconcilable.
YouTube and long-form AI video prompts (56–60)
Long-form footage has a different job. It has to survive being cut against other shots, it has to hold for several seconds without becoming boring, and it usually needs headroom for titles or a lower third. These five are written as usable coverage rather than showpieces, which is why the camera moves are conservative and the compositions leave deliberate empty space.
Wide shot, 16:9, 24mm, very slow push in over eight seconds. A single wooden desk with a closed laptop sits in a large empty room with a bare concrete floor, subject placed on the right third leaving the left half of frame empty for titles. Early morning light through tall windows camera-right, long soft shadows stretching across the floor, dust visible in the beams. Sound: room tone with a slight echo, one distant street sound, no music, no dialogue. No on-screen text, no people, constant slow camera speed, no lens flare.
The empty left half is the point. Do not let the model fill it.
Why it works: Editors need negative space and models hate it, so you have to demand it by naming which third the subject occupies. A very slow push over a stated duration gives the model a rate to hold, which prevents the drifting acceleration that makes long moves unusable. Dust in the light beams gives the frame motion even though nothing in it is moving, which is exactly what an establishing shot needs to hold for eight seconds.
Medium close-up, 16:9, 50mm, static camera at a slight high angle. Hands type steadily on a mechanical keyboard on a dark desk, continuous typing for the whole shot, no pauses. Single warm desk lamp camera-left with a cool window fill from behind, shallow depth of field, monitor glow visible but the screen itself out of frame. Sound: mechanical key clicks, a quiet fan hum, no music, no dialogue. No on-screen text, no legible screen content, no extra fingers, real-time speed.
No extra fingers is not paranoia. Keyboards are still where hands fail.
Why it works: Hands on keys remain one of the hardest things to render, so the prompt stacks the odds: a slight high angle hides the most failure-prone knuckle geometry, shallow depth softens the finger tips, and the explicit no extra fingers constraint catches what is left. Keeping the screen out of frame removes the garbled-interface problem entirely rather than trying to constrain it.
Wide shot, 16:9, 35mm, locked-off camera. Slow-moving low clouds pass behind a plain modernist rooftop parapet in the lower third of frame, the upper two thirds sky, nothing else in the shot. Overcast diffuse light, cool neutral palette, no sun disc visible. Motion is continuous and even across the whole clip. Sound: high wind, no traffic, no music, no dialogue. No on-screen text, no birds, no aircraft, no time-lapse, real-time cloud speed.
A plate is meant to be ignored. Anything eye-catching in it is a bug.
Why it works: Background plates fail when they compete with the foreground, so this one is deliberately empty of incident and heavy on sky, which is where your graphics or subject will sit. Banning time-lapse is important because models associate moving clouds with accelerated footage and will speed-ramp them unless told not to. The result is a plate that reads as a real locked-off camera left running.
Medium shot, 16:9, 85mm, static camera at eye level with the subject on the left third looking across the empty right side of frame. A man in his fifties in a grey work shirt sits on a stool, listening, occasionally blinking, no speech. Soft key from a large source camera-left at forty five degrees, subtle fill, dark unlit workshop behind him falling into shadow. Sound: workshop room tone, a distant compressor, no music, no dialogue. No on-screen text, no lip movement, no face morphing, no camera movement.
Explicitly say no lip movement or the model will invent speech.
Why it works: Interview coverage without dialogue is enormously useful and almost never prompted correctly, because any seated talking-head framing makes the model start mouthing words. Naming no lip movement alongside no dialogue closes that. Looking across the empty side of frame is the standard interview composition and gives you room for a lower third without repositioning the subject.
Close-up, 16:9, 50mm, slow lateral dolly left to right. A single brass key is placed onto a worn wooden bench by one hand and picked up by another, one continuous exchange completed inside the shot. Hard single overhead work light, strong shadow directly beneath the key, dark surroundings. Sound: metal on wood, cloth movement, no music, no dialogue. No on-screen text, no faces in frame, no extra limbs, realistic metal weight and shadow contact.
Contact shadows are what sell small-object handoffs. Ask for them by name.
Why it works: Transitions need a self-contained action, and a handoff has an unmistakable start and end. The failure mode with small objects is that they float, which is a contact-shadow problem, so the prompt names the shadow directly beneath the key. A slow lateral dolly reads as an intentional edit device rather than a camera wandering, which is what makes this cut cleanly between two sections.
Documentary and observational prompts (61–65)
Documentary style is a discipline of restraint. The camera is late, imperfect and uninvolved, the light is whatever was already there, and nobody performs for the lens. Prompts that ask for beauty here get commercials instead. These five ask for observation, and the realistic generator is the right mode for all of them.
Wide shot, 35mm, handheld with slight uncorrected drift, camera standing still among the crowd. A vendor in her sixties lifts a folded tarpaulin off a fruit stall and hooks it back, revealing stacked crates, while shoppers pass in front of the lens and briefly block the view. Early morning overcast light, no sun, damp pavement, mixed and unflattering ambient colour. Sound: overlapping voices, crates on concrete, a scooter passing, no music, no dialogue in focus. No on-screen text, no one looking at camera, no colour grading, real-time speed.
Letting a passerby block the lens is the single most documentary instruction in this list.
Why it works: Real observational footage is obstructed, and obstruction is something you have to request because models compose for clarity by default. A foreground passerby also gives the shot genuine depth layering. Specifying that nobody looks at camera prevents the eye-contact-with-the-lens tell that turns documentary footage into an advert, and asking for no colour grading keeps the ugly mixed ambient colour that real markets actually have.
Medium close-up, 50mm, handheld resting on a surface, minimal movement. A watchmaker's hands seat a tiny gear into an open movement with tweezers, one placement completed during the shot, the task clearly already underway when the clip begins. Single angled task lamp from camera-right, hard small-source light, bright specular highlights on metal, everything outside the working area dark. Sound: tweezer contact, a chair shift, faint breathing, no music, no dialogue. No on-screen text, no extra fingers, no slow motion, realistic metal reflectivity.
Beginning mid-task, not at the start of the task, is what makes it read as documentary.
Why it works: Documentary shots almost never catch the beginning of an action, so telling the model the task is already underway removes the staged, performed-for-camera quality. Hard small-source light is correct for craft work and gives the model strong specular anchors on metal, which is far more convincing than soft product lighting. Resting handheld is a real camera behaviour and reads more honest than a locked-off tripod here.
Medium wide, 35mm, static camera at eye level. A dockworker in high-visibility waterproofs stands beside a stack of mooring ropes, doing nothing in particular, shifting his weight once and looking off toward the water. Flat grey coastal daylight, no sun, wet concrete, containers stacked out of focus behind. Sound: gulls, water slapping a hull, a distant machine, no music, no dialogue. No on-screen text, no eye contact with camera, no posing, no colour grading, real-time speed.
Doing nothing in particular is a real instruction and it changes the output.
Why it works: Ask a model for a portrait and it stages one. Asking for an unperformed moment, with a single weight shift as the only action, produces the in-between beat that documentary editors actually use. The flat grey coastal light is unglamorous on purpose, and naming the absence of sun stops the golden-hour default that would immediately make this feel like a brand film.
Medium shot, 28mm, handheld with noticeable operator movement and occasional slow reframing. An empty classroom with chairs stacked on desks, late in the day, camera slowly panning across the room and pausing on a chalkboard that has been wiped but not cleaned. Low warm sun through blinds camera-left casting hard horizontal stripes across the far wall. Sound: building hum, a door closing somewhere else, no music, no dialogue. No on-screen text, no legible writing on the board, no people, real-time speed, no stabilisation.
No stabilisation is what buys you the archival texture. Do not add a gimbal.
Why it works: Overly smooth movement is the strongest signal that footage was generated, so requesting visible operator movement and no stabilisation buys authenticity for free. Blind stripes give the model a clear, high-contrast lighting structure to hold as the camera pans, which stabilises exposure across the move. Pausing on a detail mid-pan is real camera operator behaviour and something models will do if you name it.
Wide shot, 24mm, locked-off camera on a tripod, no movement at all. A bus shelter at dusk with three strangers waiting, none interacting, one checking a phone with the screen not visible to camera. Mixed light: cool blue dusk sky, warm sodium street lamp overhead, cold shelter fluorescent inside. Wet pavement reflecting all three sources. Sound: traffic, rain drip, a bus in the distance, no music, no dialogue. No on-screen text, no one looking at camera, no legible phone screen, real-time speed.
Three light temperatures in one frame is the most photographic thing you can ask for.
Why it works: Mixed colour temperature is what real night exteriors look like and what stock AI footage never has, because models default to one unified colour cast. Naming three sources with three different temperatures and a reflective ground forces the render to build a genuinely layered light map. A fully locked-off frame with unrelated people is a classic observational tableau and it is very hard for the model to get wrong.
Music video and performance prompts (66–70)
Music video is the one category where stylisation is the goal rather than a risk, but it still fails for the usual reason: asking for a vibe instead of a mechanism. Strobe, smoke, colour gels and slow motion all work well in current models when you say exactly what they do to the light. When you ask for energy, you get nothing. The camera-language piece on camera movement that makes footage feel directed pairs well with this section.
Medium shot, 50mm, slow orbit around the subject at constant speed. A singer in a black turtleneck stands still at a microphone, eyes closed, no lip movement, small head tilt as the only motion. One hard magenta gelled light from camera-right and one weak cyan rim from behind, deep black surroundings with no visible walls or floor edge. Heavy atmospheric haze making both light beams visible in the air. Sound: room tone only, no music, no dialogue. No on-screen text, no lip sync, no strobing, realistic haze behaviour, constant orbit speed.
Ask for visible beams in haze and the two-colour lighting becomes the whole set.
Why it works: Haze turns light into geometry, which gives an orbiting camera something to reveal as it moves and prevents the empty-black-void look that kills most performance prompts. Two opposed gel colours give the model an unambiguous colour map instead of an averaged purple wash. Specifying no lip movement matters because a microphone in frame will otherwise trigger mouthed singing that will never match your track.
Wide shot, 24mm, handheld at chest height moving slowly through a crowd. A packed dark venue, dozens of raised hands, bodies moving continuously, no individual face held in focus. Hard white strobe from behind the crowd firing at a steady rhythm, freezing silhouettes against smoke, complete darkness between flashes. Sound: crowd roar, low sub rumble, no music, no dialogue. No on-screen text, no faces in sharp focus, no slow motion, steady strobe rate, no epilepsy-inducing rapid flicker.
Steady strobe rate. An unspecified strobe becomes random flicker.
Why it works: Strobes work beautifully in current models because they simplify the render: each flash is a high-contrast silhouette moment and the darkness between hides everything the model would otherwise get wrong. The failure is rate, so name it as steady. Keeping every face out of sharp focus is the deliberate choice that lets a crowd of AI-generated people pass, since crowd faces are still the weakest link.
Medium close-up, 85mm, static camera. A dancer's hand strikes the surface of a shallow black reflecting pool, sending a crown of droplets upward, filmed in genuine high-speed slow motion at a constant rate. Single hard top light directly above, water lit from above only, background completely black. Sound: a single water impact stretched, low room tone, no music, no dialogue. Slow motion requested deliberately and held constant, no speed ramping, no on-screen text, realistic water droplet physics and surface tension.
This is the only prompt in the collection that asks for slow motion. Say constant, not ramped.
Why it works: Slow motion is banned in most of these prompts because models add it uninvited, but when you actually want it you have to request it explicitly and then forbid ramping, or the clip will accelerate halfway through. Water against pure black with hard top light is the classic high-speed setup and it gives the model maximum contrast on the droplets, which is where the detail lives.
Wide shot, 35mm, very slow dolly right, constant speed. An empty motel room at night with the door standing open to a lit parking lot, curtains moving slightly, a single unmade bed, nobody present. Cold blue exterior light through the open door, one warm bedside lamp inside, hard shadow shapes across the carpet. Sound: distant highway, an air conditioning unit, a moth against a bulb, no music, no dialogue. No on-screen text, no people, no signage text, realistic curtain movement, real-time speed.
An empty room with a story in it beats a performer with none.
Why it works: Absence is a legitimate music video device and it is far easier to render well than a performer, because there is no face, no lip sync and no body mechanics to fail. The prompt still gives the frame motion through the curtains and the dolly, so it does not feel like a still. Banning signage text is specific to motel settings, where the model will otherwise cover every surface in invented lettering.
Close-up, 85mm, static camera with a very slight handheld drift. A shallow puddle on dark asphalt reflecting a red and blue neon sign, the reflection rippling continuously as unseen traffic passes, sign itself never in frame. Night, wet ground, no other light source. Sound: passing tyres on wet road, distant bass through a wall, no music, no dialogue. No on-screen text, no legible letters in the reflection, no people, realistic water ripple physics, real-time speed.
Keeping the sign out of frame is how you get neon without hallucinated words.
Why it works: Neon signage is a text hallucination magnet, so the reflection-only approach gives you the colour and mood while making illegibility physically plausible. Continuous ripples driven by unseen traffic give the model an ongoing motion source that never needs to resolve into anything, which is exactly what an insert shot needs to survive being cut to a beat.
Sports and motion AI video prompts (71–75)
Sports prompts fail on physics, not on looks. Bodies at speed have weight, ground contact and follow-through, and when a prompt does not mention them the render produces gliding figures that never quite touch the surface. Every prompt below names the contact and the speed, which is the difference between athletic and floaty.
Medium shot, 85mm long lens with compressed background, static camera at track level. A sprinter drives out of the starting blocks and takes four powerful strides across frame left to right, spikes visibly gripping and flexing the track surface, full body weight visible in each ground contact. Bright overcast stadium daylight, no hard sun, blurred out-of-focus crowd behind. Sound: spikes on track, explosive breathing, a distant crowd, no music, no dialogue. Real-time speed, no slow motion, no speed ramping, realistic body weight and ground contact, no on-screen text.
Name the ground contact or the athlete will float.
Why it works: Every unconvincing AI sports clip has the same problem: the feet do not load. Naming visible grip, flex and body weight on each contact forces the model to render the compression that reads as force. Long lens compression is also standard for track and it conveniently defocuses the crowd, removing dozens of faces that would otherwise need to hold up.
Medium shot, 35mm, handheld from the baseline following the action laterally. A player drives past a defender toward the basket, plants the left foot hard, and rises, the sequence ending at the top of the jump. Indoor arena lighting from a high grid, hard shadows directly beneath the players on a polished reflective floor. Sound: sneaker squeak, a ball bounce, crowd noise, no music, no dialogue. Real-time speed, no slow motion, realistic momentum and floor contact, unbranded kit, no legible numbers or logos, no on-screen text.
End the clip at the top of the jump. Do not make the model land him.
Why it works: Landings are harder than takeoffs, so ending the described action at the apex avoids the most failure-prone frames entirely. The squeak and the reflective floor are both physics cues that reinforce contact. Requesting unbranded kit with illegible numbers heads off the garbled jersey typography that appears in almost every AI sports render.
Close-up, 50mm, camera at water surface level moving alongside the swimmer at matched speed. An open-water swimmer takes three freestyle strokes, head turning to breathe once, water breaking over the lens between strokes. Bright low morning sun from behind creating hard glitter on the chop. Sound: water rushing past the microphone, breathing, muffled underwater moments, no music, no dialogue. Real-time speed, no slow motion, realistic water displacement and splash physics, no on-screen text, no lens flare artifacts.
Water over the lens is a camera imperfection that instantly reads as real.
Why it works: Matched-speed tracking keeps the swimmer stable in frame while everything around them moves, which is a much easier render than a camera that outpaces its subject. Deliberately allowing water to break over the lens is an authenticity cue no clean render will produce on its own, and specifying displacement physics stops the swimmer from cutting through water like a knife with no wake.
Medium wide, 200mm long lens with heavy compression, static camera at road level as the group approaches. A tight group of cyclists rides directly toward camera on a mountain road, bodies rising and falling with pedal cadence, the compression stacking riders on top of each other. Hazy afternoon backlight, heat shimmer visible on the road surface. Sound: freewheel clicking, tyres on tarmac, breathing, wind, no music, no dialogue. Real-time speed, constant cadence, realistic bike geometry and wheel rotation, unbranded kit, no on-screen text, no slow motion.
Heavy compression is doing the work. It hides the bikes the model would get wrong.
Why it works: Long lens compression on an approaching group is the signature look of cycling coverage and it is also a rendering advantage, because stacked riders occlude each other and reduce the amount of visible bicycle geometry the model has to keep coherent. Naming wheel rotation and constant cadence prevents the sliding, non-rotating wheels that are the most common failure in bike footage.
Wide shot, 24mm, locked-off camera positioned below and to the side. A climber on an overhanging rock face reaches for a single hold above, body swinging outward under load, one continuous reach completed in the shot. Hard midday sun from camera-left, deep shadow in the rock, bright rock face, dark sky through the frame edge. Sound: chalk, fabric against rock, breathing, wind, no music, no dialogue. Real-time speed, no slow motion, realistic body weight and rope tension, no on-screen text, no extra limbs.
Rope tension is the physics word that makes hanging bodies believable.
Why it works: Suspended bodies are a weight problem, and the model needs to know something is pulling. Naming rope tension and body swing under load gives it that force. A locked-off wide from below is real climbing camera placement and, more usefully, keeps the camera out of the physics simulation entirely so only the athlete has to be correct.
Historical and period AI video prompts (76–80)
Period prompts go wrong when they name a decade and stop. A year is not a visual instruction, it is a lookup, and the model returns an averaged costume-drama pastiche. What actually creates a period is materials, light sources available at the time, and the capture format the era would have been recorded on. Name those three and the year takes care of itself.
Medium shot, 50mm, static camera at seated eye level. A woman in a heavy wool dress sits reading at a plain oak table, turning one page during the shot. The only light in the room comes from three candles on the table, warm and unsteady, falling off to complete darkness within two metres, faces lit from below and to the side. Rough plaster walls, no other objects. Sound: a fire in another room, the page turning, no music, no dialogue. No on-screen text, no electric light, no lens flare, realistic candle flicker and falloff.
Specify the falloff distance. That is what makes a room feel pre-electric.
Why it works: Historical interiors read as fake when they are evenly lit, because even lighting requires technology the period did not have. Naming the light source, the direction, and the exact distance at which it dies gives the model the physical constraint that produces the period look. Underlighting from a table is a genuine artefact of candle placement and it is not something the model will choose on its own.
Medium wide, 25mm, handheld with visible operator instability, shot on 16mm black and white film with pronounced grain, slight gate weave and occasional dust. Workers leave a factory gate in a steady stream, some glancing toward the camera, one raising a hand. Flat overcast daylight, high contrast black and white with crushed shadows. Sound: crowd footsteps, industrial ambience, no music, no dialogue. No on-screen text, no title cards, no colour, no digital sharpness, real-time speed.
Ask for the format, not the year. Grain and gate weave are the period.
Why it works: Capture format carries more period information than costume does. Naming 16mm, grain, gate weave and dust tells the model exactly which imperfections to synthesise, and those imperfections are what the eye reads as archival. Banning digital sharpness matters because models default to a clean modern image with a grain overlay on top, which fools nobody.
Wide shot, 35mm, slow handheld drift. A cobbled street at dawn with horse-drawn cart tracks in wet mud, timber-framed buildings with small irregular window panes, smoke rising from two chimneys. One figure in a long coat walks away from camera into the mist. Cold blue pre-sunrise light with no direct sun, everything low contrast and desaturated by haze. Sound: distant hooves, a shutter opening, birds, no music, no dialogue. No on-screen text, no signage, no modern materials, no glass storefronts, no tarmac, real-time speed.
Banning the anachronisms explicitly works better than naming the century.
Why it works: Models leak modern materials into period scenes constantly: smooth tarmac, plate glass, painted signage. Listing the specific things that must not appear is far more effective than trusting a date to exclude them. Dawn haze is also a practical choice, since low contrast and reduced detail give the render fewer opportunities to place a modern object in sharp focus.
Medium close-up, 50mm, static camera. A cooper shapes a barrel stave with a hand plane on a wooden bench, three continuous strokes, shavings curling and falling. Light comes from one small unglazed window camera-left, hard-edged and dusty, the rest of the workshop dark. Iron tools on a rack, wood shavings underfoot. Sound: plane on wood, shavings falling, a distant bell, no music, no dialogue. No on-screen text, no modern tools, no electric light, no extra limbs, realistic wood and shaving physics.
One unglazed window is a period lighting rig and it produces beautiful hard light.
Why it works: Craft process shots suit historical prompting because the action is repetitive and physical, which models handle well, and the tools are simple enough to render correctly. A single small window gives hard directional light with visible dust, which is both period-accurate and technically flattering. The three-stroke instruction bounds the action so it completes within the clip.
Wide shot, 28mm, static camera on a tripod at platform level. A crowded railway platform under a glazed iron roof, steam drifting across the frame from a locomotive out of shot, people standing with cases, nobody speaking. Grey daylight filtered through dirty glass overhead, strong shafts where panes are broken, steam catching the light. Sound: steam venting, footsteps on stone, a whistle, no music, no dialogue. No on-screen text, no legible signage or destination boards, no modern clothing, real-time speed, realistic steam behaviour.
Steam is a gift to period prompts. It hides detail and it dates the scene at once.
Why it works: Volumetric atmosphere both establishes the era and conceals the background detail most likely to break. Broken panes give the model a reason to render light shafts, which makes the steam visible and structures the whole frame. Requesting illegible destination boards is the right call, since a period station without signage looks wrong but rendered signage will be gibberish.
Automotive AI video prompts (81–85)
Cars are the single most demanding subject in AI video, because a car is a rigid body with reflective panels, rotating wheels and precise proportions, and viewers know instantly when any of the three is wrong. The fixes are consistent: name the wheel rotation, name the reflections, keep the badge off screen, and prefer long lenses that reduce visible panel geometry.
Medium wide, 50mm, camera tracking alongside the vehicle at matched speed from a parallel road. A dark grey unbranded sedan drives at constant speed along a coastal road, wheels rotating correctly with visible spoke motion, suspension reacting to small road imperfections. Late afternoon sun low from camera-right, long hard reflections sliding along the door panels as the car moves, sea out of focus beyond. Sound: tyre roar, wind, engine at steady load, no music, no dialogue. Real-time speed, no slow motion, no speed ramping, correct wheel rotation direction, realistic reflection movement, no badges, no logos, no on-screen text.
Sliding reflections are the tell. Ask for them and the panels stop looking like plastic.
Why it works: A car body is defined by what moves across it, so naming reflections that slide along the panels as the car travels gives the model the single most important motion cue. Correct wheel rotation direction is worth stating outright because reverse-spinning wheels remain a frequent artefact. Suspension reacting to imperfections adds the mass that separates a real vehicle from a floating model.
Medium close-up, 35mm, camera mounted on the dashboard looking back at the driver, slight constant vibration from the road. A driver in her forties holds the wheel, makes one small correction, eyes on the road, not on camera. Overcast daylight through the windscreen, cool flat light on her face, reflections of passing trees moving across the side window. Sound: road noise, indicator tick, ventilation hum, no music, no dialogue. Real-time speed, constant road vibration, realistic hand position and wheel grip, no extra fingers, no on-screen text, no phone in frame.
Dashboard vibration is a two-word instruction that makes the whole shot believable.
Why it works: Vehicle interiors are static frames with a moving world outside, so all the realism has to come from vibration and passing reflections. Naming both gives the model its motion budget. A single small steering correction is enough action for the clip and avoids the oversteering, video-game wheel movement models produce when told simply to drive.
Wide shot, 35mm, static camera at kerb level as the vehicle passes left to right. An unbranded dark hatchback drives past on a wet city street at night, headlights throwing a moving pool of light across the road surface, tail lights leaving a red wash on the asphalt behind. Mixed sodium and white street lighting overhead, heavy reflections on the wet ground, no other traffic. Sound: tyres through standing water, engine passing, distant city, no music, no dialogue. Real-time speed, realistic light pool movement and water spray, correct wheel rotation, no badges, no legible plates, no on-screen text.
Light pools moving with the car are the physics detail that makes night driving read.
Why it works: Night automotive shots live or die on whether the car's own lights affect the environment. Naming the moving headlight pool and the tail light wash forces the model to link the vehicle to the road rather than compositing it on top. Illegible plates is a small but necessary constraint, since number plates are guaranteed gibberish.
Close-up, 85mm, static camera at ground level. A single tyre rolls slowly into frame across loose gravel and stops, individual stones displacing and settling under the load, tread visibly compressing. Hard low sun from camera-left, long shadows in the gravel texture, dust raised slightly and drifting. Sound: gravel crunch under load, a settling stone, no music, no dialogue. Real-time speed, realistic tyre deformation and gravel displacement, no slow motion, no branding on the sidewall, no on-screen text.
Tread compression under load is the detail almost no AI car shot has.
Why it works: Insert shots are the easiest automotive wins because they crop out the proportions that models get wrong and focus on materials, which they handle well. Naming deformation and displacement makes the tyre interact with the ground instead of resting on it. Sidewall lettering is guaranteed to hallucinate, so it is banned.
Medium wide, 85mm, very slow orbit at constant speed around a stationary vehicle. An unbranded matte grey coupe sits on a seamless dark grey studio floor, no people, no props. One very large soft overhead source producing a single continuous soft highlight running the length of the roof and shoulder line, everything else falling to near black, no visible light stands or crew reflections. Sound: quiet studio room tone, no music, no dialogue. Constant orbit speed, realistic reflection travel across the panels, no badges, no logos, no on-screen text, no lens flare.
One continuous highlight along the shoulder line is how real car photographers light a body.
Why it works: Automotive studio lighting is a single enormous source, and describing that source and the highlight it creates gives the model an accurate physical setup instead of scattered speculars. Explicitly excluding crew reflections matters because a reflective body in a studio will otherwise render distorted figures in the paint. The slow constant orbit lets the highlight travel, which is what shows the shape.
Architecture and interior prompts (86–90)
Architecture is a light problem disguised as a geometry problem. The building is usually rendered fine; what fails is that every surface is lit identically, so the space has no depth. These five each name a single dominant source and describe what it does to the surfaces, which is what produces the volume that architectural footage needs.
Wide shot, 24mm, very slow push forward at constant speed. An empty concrete gallery space with a single rectangular skylight above, the light falling as one sharp bright shape on the floor and gradually falling off into dim grey corners. Board-marked concrete walls with visible timber grain texture, polished floor faintly reflecting the light shape. No people, no furniture. Sound: deep room reverb, a distant door, no music, no dialogue. No on-screen text, no people, constant camera speed, realistic light falloff, no additional light sources.
No additional light sources is the clause that creates the drama.
Why it works: Models will add fill light to any dark corner unless told not to, which flattens architectural interiors into estate agent footage. Forbidding extra sources preserves the falloff that gives the space volume. Naming board-marked concrete gives the surfaces a real texture to render, and the faint floor reflection ties the light shape to the material.
Wide shot, 35mm, static camera, no movement. A brick and glass facade catching very low morning sun from an extreme side angle, every course of brickwork throwing its own small shadow, the glass reflecting a cold blue sky. Deep shadow across the lower third where an adjacent building blocks the sun. Sound: light traffic, birds, wind, no music, no dialogue. No on-screen text, no people, no signage, realistic material reflectivity, no lens flare, no colour grading.
Raking light turns flat brick into texture. It is the cheapest architectural upgrade there is.
Why it works: Extreme side lighting makes surface relief visible, and naming the per-course shadow tells the model to render brick as geometry rather than as a texture map. The blocked lower third introduces a second tonal zone that stops the facade reading as a flat elevation. No colour grading keeps the real contrast between warm brick and cold sky.
Medium wide, 28mm, smooth forward tracking shot at walking pace through a doorway into a living space. A furnished modern apartment lit only by daylight from a large window on the far wall, brightness dropping sharply as the camera passes through the darker doorway into the lit room. No artificial lights on. Wood floor, linen sofa, no clutter, no people. Sound: quiet interior room tone, distant street, no music, no dialogue. No on-screen text, no people, no televisions on, realistic exposure change through the doorway, real-time speed.
The exposure change through the doorway is what makes a walkthrough feel filmed.
Why it works: Real cameras change exposure as they move between zones of different brightness, and asking for that shift makes a walkthrough feel captured rather than rendered. Turning off every artificial source forces the model to commit to one light direction. Keeping screens off removes another guaranteed source of garbled content.
Close-up to medium, 50mm, static camera looking straight down a spiral staircase from the top. A tight helical stair descending out of view, handrail curve unbroken, treads receding evenly, no people. Soft daylight from a window several floors down giving a gradient from dark at the top to bright at the bottom of the shot. Sound: faint building reverb, no music, no dialogue. No on-screen text, no people, no camera movement, consistent stair geometry with even tread spacing, no warping.
Even tread spacing and no warping. Spiral stairs are where geometry breaks.
Why it works: Repeating geometry is the classic failure case, so the constraints here target it directly: unbroken handrail, even treads, no warping. The vertical light gradient gives the shot direction and depth without any camera movement, and a locked-off frame means the geometry only has to be right once rather than through a move.
Wide shot, 35mm, locked-off camera. A modern house exterior at blue hour with warm interior lights visible through large windows, one additional room light switching on partway through the shot. Cold blue ambient sky, warm tungsten pools spilling onto a terrace, no direct sun, silhouetted planting in the foreground. Sound: evening insects, a distant road, no music, no dialogue. No on-screen text, no people visible inside, realistic light spill onto exterior surfaces, real-time speed, no time-lapse.
One light switching on is enough narrative for an architectural shot.
Why it works: Blue hour is the standard architectural exterior for a reason: interior warm and exterior cool balance in exposure, which gives you both at once. Adding a single light change provides a story beat without a person, and naming the spill onto the terrace connects the interior source to the exterior surfaces. Banning time-lapse keeps the sky from racing.
Dialogue and lip-sync AI video prompts (91–95)
Dialogue is the one place where the prompt format genuinely changes. The exact line goes in quotes, there is one speaker, and the line has to be short enough to finish inside the clip. Everything else is normal prompting. If you have ever had a character mouth words that never matched your audio, it is almost always because the line was implied rather than quoted.
Medium close-up, 50mm, static camera at eye level. A woman in her thirties in a plain navy jumper looks directly into the lens and says, in a calm even tone: "We tried this three times before it worked." She finishes the line and holds still. Soft key light from a large source camera-left, gentle fill, plain warm grey wall behind, shallow depth of field. Sound: quiet room tone, her voice clearly foregrounded, no music. Lip movement matches the quoted line exactly, one speaker only, no off-screen voices, no on-screen text, no face morphing.
The line is short on purpose. Long lines drift out of sync before they end.
Why it works: Quoting the exact line is what activates dialogue generation instead of generic mouthing, and keeping it under about ten words means it finishes comfortably within the clip rather than being truncated or rushed. Naming one speaker and banning off-screen voices prevents the phantom second voice that appears when the model senses a conversation. Static camera and soft key keep the face stable, which is where sync artefacts show first.
Medium two-shot, 35mm, static camera at seated eye level. Two men sit across a small table. The older one says: "You already know the answer." The younger one replies: "I wanted to hear you say it." Only one speaks at a time, no overlap, both lines complete within the shot. Warm practical lamp between them, soft falloff, dark bar interior behind. Sound: quiet bar ambience, both voices clear, no music. Lip movement matches each quoted line for the correct speaker, no crosstalk, no on-screen text, no face morphing, no extra limbs.
Say no overlap. Models will happily have both characters talk at once.
Why it works: Two-handers work when the turn-taking is explicit. Assigning each quoted line to a described character and forbidding overlap gives the model an unambiguous timeline. Keeping both faces in one static frame avoids cutting, which is where character identity drifts, and the single practical lamp gives both faces a shared light direction so they read as being in the same room.
Medium shot, 50mm, static camera. A kitchen table with a half-finished meal, hands visible resting on the table, the speaker's face out of frame above the top edge. A man's voice, close and unhurried, says: "It was never about the money." Nothing else moves. Late evening light from a window camera-right, warm and low. Sound: room tone, cutlery settling, his voice close-miked, no music. No lip sync required, no face in frame, no second voice, no on-screen text, real-time speed.
No face means no sync risk. This is the safest dialogue shot you can generate.
Why it works: Off-camera dialogue removes lip sync from the equation entirely while keeping the emotional content, which makes it the highest-reliability dialogue prompt in the set. The prompt still has to state that no lip sync is required, or the model will crop a mouth back into frame. Close-miked voice with quiet room tone is the audio direction that makes it feel intimate rather than dubbed.
Medium close-up, 85mm, static camera with slight handheld drift. A man in his forties stands at a window holding a phone to his ear and says: "I am not coming back tonight." He listens, jaw tightening, no further speech. Cold overcast daylight from the window in front of him, his face lit flatly, the room behind him dark. Sound: room tone, faint unintelligible voice from the phone, his voice clear, no music. Lip movement matches only the quoted line, silence after, no second visible speaker, no on-screen text, no face morphing.
The listening half is the performance. Tell the model the speech has ended.
Why it works: Half a conversation is a strong dramatic device and it halves the sync workload. The critical instruction is that the lip movement stops after the quoted line, otherwise the model keeps him talking through the listening beat. An unintelligible voice from the phone gives the scene a counterpart without generating a second syncable performance.
Medium shot, 35mm, static camera at chest height. A woman in an apron stands at a workbench, holds up a small clay bowl, looks into the lens and says: "This is the stage where most people rush." She then sets the bowl down. Even soft daylight from a large window camera-left, bright clean workshop, shallow depth behind. Sound: workshop room tone, her voice clear and close, no music. Lip movement matches the quoted line exactly, one speaker, action and line complete within the shot, no on-screen text, no extra limbs, real-time speed.
Pair a short line with a small physical action and the whole clip feels directed.
Why it works: Combining a quoted line with a simple prop action gives the shot a beginning and an end, which prevents the awkward held frame that follows most dialogue renders. Naming that both the line and the action complete inside the shot bounds the timing. Even soft daylight is the right choice for instructional content and it keeps the face stable through the speech.
Image-to-video prompts (96–100)
Image-to-video follows one rule that overturns everything above: do not describe appearance. The still already carries the subject, the colour and the composition, and re-describing them invites the model to redesign your frame. You are writing a motion brief, not a shot description. These five are templates you fill with your own image, and they are deliberately short.
Animate the provided image. Motion only: the subject breathes slowly, blinks twice at natural intervals, and turns their head approximately ten degrees to their left, then stops. Hair moves slightly with the head turn. The camera performs a very slow push in at constant speed. Do not change the subject's face, clothing, hairstyle, colour palette, or background. Do not add or remove any objects. No on-screen text, no slow motion, no relighting, no style change, real-time speed, preserve the original composition and framing.
Preserve the original composition is the most important sentence in image-to-video.
Why it works: The dominant failure in image-to-video is drift: the model gradually rebuilds the subject into its own preferred version. Explicitly forbidding changes to face, clothing, palette and background anchors it to the source. Ten degrees is a deliberately small, specified rotation, because unquantified head turns overshoot and reveal geometry that was never in the original frame.
Animate the provided image. Motion only: clouds drift slowly left to right at a constant rate, water surface ripples continuously, vegetation moves gently in a steady breeze. No other movement. The camera remains completely locked off. Do not change the light direction, colour grade, time of day, or any solid geometry in the frame. No new objects, no birds, no aircraft, no people. No on-screen text, no time-lapse, real-time speed, preserve the original framing exactly.
Name every element that should move and forbid everything else.
Why it works: Landscapes have a small set of legitimately movable elements, and listing them exhaustively while banning all other motion prevents the model from animating rocks or shifting a horizon. Locking the camera removes parallax, which is where still-derived geometry falls apart, since the model has no information about what lies behind anything.
Animate the provided image. Motion only: the object rotates slowly around its vertical axis by approximately thirty degrees at a constant rate, with the highlight travelling across its surface as it turns. The camera stays static. Do not change the object's shape, material, colour, label placement, or the background. Do not add reflections of people or equipment. No on-screen text, no new branding, no logos, no style change, real-time speed, preserve the original lighting setup.
Thirty degrees, not a full turn. Beyond that the model has to invent the back.
Why it works: A still contains no information about the unseen sides of an object, so a small rotation stays within what can be inferred while a full turntable forces invention. Naming the travelling highlight keeps the lighting attached to the studio setup in the source image. Forbidding new branding is essential, since a rotating product is the exact condition under which logos hallucinate.
Animate the provided image. Motion only: light shifts very slightly as if a cloud passes, a curtain moves gently, dust drifts in the light beam. The camera performs an extremely slow push in at constant speed. Do not alter the architecture, materials, furniture positions, or colour palette. Do not add people, animals, or new light sources. No on-screen text, no relighting, no style change, real-time speed, preserve the original perspective and vertical lines.
Preserve vertical lines. A push in on architecture will otherwise start to lean.
Why it works: Architectural stills have strict perspective, and any camera move risks converging verticals that were straight in the source. Naming them protects the geometry. The three permitted motions are all atmospheric rather than structural, which is exactly the kind of movement a still can support without new information.
Animate the provided image. Motion only: the subject completes one small repeatable action, raising a cup toward the mouth and lowering it again, ending in the exact starting pose so the clip can loop. Breathing continues throughout. The camera stays locked off. Do not change the face, hands, clothing, background, or lighting. Do not add objects that are not already in the frame. No on-screen text, no extra limbs, no face morphing, no style change, real-time speed, preserve the original framing.
Ending in the starting pose is what makes an image-to-video clip loop cleanly.
Why it works: Returning to the source pose means the last frame matches the first, which is both a loop and a built-in accuracy check: if the final frame no longer resembles your original image, the model drifted. Constraining the action to objects already present in the frame prevents the invented props that appear when a model is asked for an action without a stated prop.
How to modify any prompt on this page without breaking it
Every prompt here has a load-bearing half and a swappable half. The swappable half is the subject and the environment. The load-bearing half is the camera clause, the light clause, and the constraint clause. If you replace the subject and keep the rest, the render usually holds. If you rewrite the camera and light clauses in your own words, quality drops, not because your words are worse but because you have probably collapsed two decisions into one adjective.
- Change the subject first, alone, and render once. This tells you whether the frame still reads.
- Change the environment second. Keep the two concrete background details, just make them yours.
- Only then touch camera or light, and change one of them per render, never both.
- Never delete the constraint line. It costs you nothing and it is the cheapest quality insurance in the whole prompt.
- Save every prompt that works next to its output file. A prompt you cannot find again is a prompt you did not really write.
That discipline matters more when you are cutting several shots together. Continuity across clips is its own skill, and the two pieces worth reading next are character consistency across shots and camera language that makes footage feel directed. If your output is vertical, vertical video that holds attention covers the framing changes that matter at 9:16.
A short note on models and where these were tested
These prompts were written against current-generation video models with native audio generation, which is the behaviour documented in Google's Veo model family and mirrored by most competing systems in 2026. If you want the primary source on how these models interpret camera and audio instructions, Google DeepMind publishes its own guidance at deepmind.google/models/veo and the developer-facing prompting notes live in the Gemini API video documentation. Nothing on this page depends on a specific vendor, but reading a first-party prompting spec once will change how you write forever.
Inside ReelVision AI, the same prompt routes to different engines depending on the mode you pick, which is why the ultra realistic prompts in this collection read differently from the cinematic ones. If you want a fast way to see that difference, run prompt 6 through the realistic generator and prompt 1 through the cinematic generator back to back. The gap between the two outputs is the whole argument for writing per-mode prompts instead of one generic prompt.
Run these prompts instead of reading about them
Every prompt on this page was written to be pasted straight into ReelVision AI. Pick one, change the subject, and watch what a decidable prompt actually produces.
Start generatingFrequently asked
The ReelVision AI editorial team writes about cinematic AI video production, prompt craft, and the workflows behind believable footage.
Try this in ReelVision AI
Everything in this article is one prompt away. Open the studio and render it with locked characters, native audio, and real cinematic motion.
Start CreatingCinematic AI craft, weekly
One email. Prompt breakdowns, camera language, and workflows that make AI footage look directed.
No spam. Unsubscribe in one click.
Related articles
TrendingThe Complete AI Video Prompt Guide (2026 Edition)
The seven-part framework we use to write AI video prompts that behave predictably, plus tested examples for cinematic, ultra realistic, anime, commercial, social, and YouTube Shorts work.
TrendingThe Anatomy of a Cinematic AI Video Prompt
Most AI clips fail for one reason: the prompt described a picture instead of a shot. Here is the six-part structure working directors use.

Character Consistency: How to Keep the Same Face Across Every Shot
Nothing breaks an AI edit faster than a lead actor whose face changes between cuts. Here is how to lock identity across an entire sequence.