← Blockshot

How to control camera movement in AI video

Video models render subjects well and camera direction badly. That asymmetry has a cause, and once you see it the workaround is obvious.

Why prompts lose the camera

A model learns from captioned video. Captions describe what was in the shot reliably and how the shot was made almost never. "A red car on a wet street" is an accurate, repeatable label. "Sweeping cinematic shot" is applied to cranes, orbits, drone passes and steady push-ins alike.

So the model's idea of a subject is sharp and its idea of camera motion is an average of everything anyone ever called cinematic. Ask for a red car and you get one. Ask for "slow dolly in, then whip pan left" and you get motion that is plausibly cinematic and rarely the motion you named.

This gets worse the more precise you are. Timing — when the pan starts relative to the action — has essentially no vocabulary in captions at all. Two moves in sequence is usually beyond what a prompt can hold.

What helps, in order of how much

  1. Separate the framing from the movement. They are two decisions and prompts routinely fuse them. "Wide, low, static" is a framing. "Push in" is a movement. A shot needs both stated, and one of them is allowed to be "none".
  2. Ask whether it needs to move at all. A locked-off frame is a real choice, not a failure to specify one. Most sequences hold better with fewer moving shots than people write into prompts, and a static shot is the one thing a model reproduces reliably.
  3. Name the move, do not evoke it. Use the actual terms — orbit, crane down, over the shoulder — rather than adjectives like sweeping or dynamic. Precise words at least narrow the average.
  4. Say what changes across the shot. "Starts wide on the doorway, ends tight on her hands" constrains a result. "Dramatic camera work" cannot.
  5. Stop describing and start showing. Everything above is damage control on a lossy channel. The channel is the problem.

Showing instead of describing

Give the model a reference clip and the ambiguity disappears. Not because the model understands cinematography better, but because it is no longer being asked to. The motion is in the pixels: frame by frame, the reference already answers where the camera is, where it is pointing and how fast it is getting there.

The clip does not need to look like the finished shot. It needs the right camera path, the right timing and objects in roughly the right places at the right size. Grey boxes standing in for a car, a doorway and two people are enough — this is exactly what previsualization has always been for, and the reason film productions have done it for decades before any of this existed.

How to do it without a 3D department

  1. Describe the scene once, in plain language — what is in it, who is doing what, and how it should be covered.
  2. Let the layout be generated. Blockshot places blocky stand-ins and proposes a shot list with named camera moves against a shared action clock.
  3. Fix what is wrong by hand. Move objects, aim the camera, draw a path, change the lens, retime a cut. This is the part a prompt cannot do at all.
  4. Export the clip and pass it to your video model as a reference or control input, alongside a prompt that now only has to carry look and content — the two things prompts are good at.

Rendering and encoding happen in the browser, so the scene never leaves the machine it was built on.

Frequently asked

Why does my AI video ignore the camera movement in my prompt?

Because camera direction is weakly represented in the captions models train on. Subject descriptions are labelled consistently; camera moves are labelled with whatever the caption writer felt. The model averages that ambiguity rather than following your instruction.

Do longer, more detailed prompts fix camera control?

Only slightly, and detail about timing or sequences of moves tends to make things worse rather than better. Precision in a channel that does not carry precision produces confident-looking noise.

What is a reference clip?

A short video handed to a generation model to show motion rather than describe it. It can be rough — grey placeholder shapes are fine — because what is being communicated is the camera path and timing, not the appearance.

Do I need to know 3D software to make one?

No. Blockshot generates the layout and the camera moves from a plain-language description, and everything is edited in the browser by dragging.

Does this work with Sora, Veo, Runway or Kling?

Any model that accepts a video input as a reference or control signal can use one. The clip is an ordinary video file, so it is not tied to a particular provider.

Blockshot turns a description into a blockout with real camera moves, and exports the clip. Try it →