Camera movement prompts for Sora, Veo, Runway and Kling
A camera movement prompt names the shot size, one camera move and the subject's action, in film terms. OpenAI, Google, Runway and Kling each publish a prompting guide for their video model. The four guides agree on one main camera move per shot, and differ in how they want the prompt laid out.
Every model-specific rule on this page comes from that model's official documentation, checked on 12 September 2026 and linked under Sources. Blockshot has not run controlled tests of these models. The worked examples follow each guide's recommended structure; they are not measured results.
Rules the four guides share
- One main camera move per shot. OpenAI's Sora 2 guide says each shot should have one clear camera move and one clear subject action. Kling's camera control guide says to keep one main camera move per shot.
- Film terms, not adjectives. All four guides use named moves such as dolly, push-in, pan, tilt, tracking shot and crane shot.
- Speed words. The guides' own examples qualify moves with words such as "slow", "gentle" and "stable".
- A start and an end. Kling's guide recommends stating the opening view, the move, and what the move reveals. Runway's guide recommends describing what comes into view at each phase of a move.
- Short clips. OpenAI's Sora 2 guide says the model follows instructions more reliably in shorter clips.
The four models at a glance
| Model | Prompt structure | Camera terms named | Reference inputs |
|---|---|---|---|
| Sora 2 | Scene description, then a cinematography section with camera and lens, then action beats, then dialogue | Shot size and angle, push-in, dolly-in, tilt, handheld | Images; separate 2–4 second videos to create reusable characters |
| Veo 3.1 | Cinematography + subject + action + context + style and ambiance | Dolly shot, tracking shot, crane shot, aerial view, slow pan, POV shot | Image to video, reference images, first and last frame |
| Runway Gen-4.5 | What should happen, including what the move reveals | A dedicated camera terms guide in Runway's help centre | Text or an input image; Aleph edits an input video |
| Kling VIDEO 3.0 | Opening view, then the camera move, then what it reveals | Push in, pull back, pan, tilt, track forward, orbit slowly, static camera | Images, videos and elements |
One shot, written for each model
The shot: a detective stands in the doorway of a dark warehouse at night, and the camera slowly pushes in on her while she steps inside. Each prompt follows the structure its vendor's guide recommends.
Sora 2
A detective in a long coat stands in the open doorway of a dark warehouse at night, rain falling behind her. Camera: medium-wide shot, slow dolly-in from eye level. Lens: 35 mm; shallow depth of field. Action: she takes two steps across the threshold, pauses, and looks toward the crates at the back.
Veo 3.1
Slow dolly shot pushing in at eye level, a detective in a long coat, stepping through a doorway and stopping to look ahead, inside a dark warehouse at night with rain falling outside, cinematic, moody, cool blue light.
Runway Gen-4.5
The camera slowly dollies in toward a detective standing in a warehouse doorway at night. As the camera moves closer, crates stacked at the back of the warehouse come into view behind her. She steps forward once and stops.
Kling VIDEO 3.0
Start with a wide view of a detective in a warehouse doorway at night, then a slow push-in toward her to reveal the rain on her coat, stable camera, cinematic lighting.
Sora 2
OpenAI's Sora 2 guide asks for one clear camera move and one clear subject action per shot. The guide names framing with phrases such as "wide establishing shot, eye level" and "medium close-up shot, slight angle from behind". The guide's examples put camera and lens on their own lines, for example "Camera: medium close-up, slow push-in with gentle parallax from hanging tools" and "Lens: 35 mm virtual lens; shallow depth of field". The guide recommends describing action as countable beats, such as "Actor takes four steps to the window, pauses, and pulls the curtain in the final second". The guide says movement is often the hardest part to get right, so keep it simple. Sora 2 accepts JPEG, PNG and WebP images as visual references. Sora 2's character feature accepts 2–4 second MP4 videos to create reusable characters.
Veo 3.1
Google's Veo 3.1 guide recommends the structure cinematography + subject + action + context + style and ambiance. The guide names the camera movements dolly shot, tracking shot, crane shot, aerial view, slow pan and POV shot. The guide names the compositions wide shot, close-up, extreme close-up, low angle and two-shot. The guide names the lens and focus terms shallow depth of field, wide-angle lens, soft focus, macro lens and deep focus. One of the guide's examples is "Crane shot starting low on a lone hiker and ascending high above, revealing they are standing on the edge of a colossal, mist-filled canyon at sunrise". Veo 3.1 makes clips of 4, 6 or 8 seconds, at 720p or 1080p, in 16:9 or 9:16. Veo 3.1 accepts a start image, reference images, and a first and last frame to transition between.
Runway Gen-4.5
Runway's Gen-4 guide says prompts should describe what should happen, not what should be avoided. Runway's guide warns that re-describing elements of an input image in high detail can reduce motion. Runway's camera terms guide says combining camera terms is encouraged with Gen-4.5 text to video. Runway's camera terms guide recommends describing what is in view, or what gets revealed, at each phase of a move through the frame. Runway's camera terms guide notes that static shots can be hard to get, most notably establishing and wide landscape shots. Runway's Aleph model edits and transforms an input video with a text prompt.
Kling VIDEO 3.0
Kling's camera control guide recommends the camera terms push in, pull back, pan left, pan right, tilt up, tilt down, track forward, orbit slowly and static camera. The guide says to keep one main camera move per shot. The guide recommends the pattern "Start with [close/wide view], then [camera movement] to reveal [final view]". One of the guide's examples is "Slow push-in toward a luxury skincare bottle on a marble table, morning sunlight, soft shadows, realistic reflections, stable camera, 4K cinematic detail." Kling VIDEO 3.0 accepts images, videos and elements as references. Kling's VIDEO 3.0 Omni guide describes video references as 3–8 second clips of a single character, used to keep that character consistent. Kling VIDEO 3.0 Omni generates up to 15 seconds in one generation, across multiple shots.
Where a blockout helps
The official guides checked for this page do not describe copying a camera path from a reference video. A blockout still gives Sora 2, Veo 3.1, Runway and Kling an exact starting frame. Veo 3.1's first and last frame input takes a start image and an end image, which a blockout can render from the first and last moment of the move. A blockout also fixes the numbers a prompt needs: shot size, camera height, focal length and how long the move takes. Blockshot writes a prompt for each shot from those numbers, and exports each shot's first frame or any frame as a PNG.
The words for each move are defined in Camera moves, defined. The reason text loses camera direction is in How to control camera movement in AI video.
Frequently asked
How do I get Sora to follow camera movement?
Give each Sora 2 shot one clear camera move and one clear subject action. Name the shot size, angle and move in film terms, such as "medium-wide shot, slow dolly-in from eye level". OpenAI's guide says Sora 2 follows instructions more reliably in shorter clips.
What camera movements does Veo 3 understand?
Google's Veo 3.1 prompting guide names dolly shot, tracking shot, crane shot, aerial view, slow pan and POV shot. The guide recommends putting the camera work first, before the subject, the action, the context and the style.
How do I control camera movement in Runway?
Describe what should happen in the shot, including the camera move and what it reveals. Runway's guide says prompts should describe what should happen rather than what to avoid. Runway's guide warns that re-describing an input image in detail can reduce motion.
What are the best Kling camera movement prompts?
Kling's guide recommends one main camera move per shot, named as push in, pull back, pan left, pan right, tilt up, tilt down, track forward, orbit slowly or static camera. Kling's guide recommends stating the opening view, the move, and what the move reveals.
Why does my AI video ignore the camera movement in my prompt?
A video model follows camera direction less reliably when a prompt asks for several moves or describes the move vaguely. OpenAI's and Kling's guides both recommend one main camera move per shot. OpenAI's Sora 2 guide says movement is often the hardest part of a prompt to get right.
Sources
- OpenAI: Sora 2 Prompting Guide
- Google Cloud: Ultimate prompting guide for Veo 3.1
- Runway: Gen-4 Video Prompting Guide
- Runway: Camera Terms, Prompts, & Examples
- Kling AI: Camera Control Guide, 13 August 2026
- Kling AI: VIDEO 3.0 Omni user guide
Blockshot blocks out the shot in 3D, writes the prompt from the real framing and lens, and exports the frames and the clip. See how it works →