← Blog

Why AI video models keep getting your camera moves wrong

· 6 min read

The lighting is gorgeous, the character is right, and the camera does something you never asked for. It is not a skill problem, and it is not only a model problem. It is a language problem, and film solved it long before AI video existed.

Anyone who has spent an evening with a video model knows the pattern. You wrote "slow dolly in toward her face" and got a zoom. You asked for the camera to orbit the car and it drifted sideways. You regenerated six times, spent the credits, and settled.

A prompt carries the look, not the move

A text prompt is very good at describing what is in a frame: subject, setting, mood, lens feel, colour. It is very bad at describing motion through space over time. A simple camera move involves:

"Dolly in" compresses all of that into two words, and the model fills in the rest from millions of clips where dolly, push, zoom and move closer are used interchangeably. You get an average of what people meant, not what you meant.

It gets worse once something in the scene moves too. "The camera follows him down the corridor, stops at the doorway and pans as he crosses the room" describes two moving things and the relationship between them. That is where text prompts break down almost completely.

Film already had the answer: previs

Big productions do not describe complex shots in words either. They use previsualization: a rough 3D version of the scene, with grey placeholder shapes for people, vehicles and sets, and a virtual camera doing exactly the move the director wants.

Previs is not meant to look good. It is meant to be unambiguous. When a stunt coordinator, a camera operator and a VFX supervisor watch the same previs clip, nobody argues about what "dolly in" means. They can see it.

AI video generation brings that need back. Most current video models accept a reference video, a motion input or first and last frames. Hand a model a clip of the camera move and you are no longer asking it to interpret language. You are showing it the geometry.

What a good reference clip contains

  1. Correct spatial relationships. Where the subject is relative to the set and to the camera.
  2. Correct scale. A person 1.8 m tall, a doorway 2.1 m high, a car 4.4 m long. Wrong scale gives wrong lens behaviour.
  3. The real camera path, with its speed and easing.
  4. The subject's motion, timed against the camera, so "stops at the doorway as he walks past" happens at the right moment.
  5. The right aspect ratio for where the video is going: wide, square or vertical.

Blocky grey shapes are enough, and arguably better than detail: the model takes composition and motion from the reference and style from your prompt, without confusing the two.

The workflow that fixes it

This works with any model that accepts reference motion.

  1. Block the scene before you write the prompt. Lay out the space: walls, doorways, furniture, the path your character walks.
  2. Direct the camera in 3D, not in words. Place the camera, set where it starts and ends, and play it back. A move that feels wrong in grey boxes will feel wrong in the render, and fixing it here costs nothing.
  3. Split the edit into shots on one timeline. Keep the action on a single clock and treat each shot as a window onto it, so the cuts stay continuous.
  4. Export one reference clip per shot, at the resolution and aspect ratio of the target platform.
  5. Write the prompt for the look only. Lighting, wardrobe, texture, mood. The camera and blocking come from the clip, so prompts get shorter and results more consistent.
  6. Use structure passes where the model supports them. Depth, normals and line-art renders of the same blockout are an even stronger signal for structure-conditioned models.

For the wording that still belongs in the prompt, see camera movement prompts for Sora, Veo, Runway and Kling.

Where Blockshot fits

We built Blockshot because we kept hitting exactly this wall, and traditional previs tools are heavy 3D packages made for studios, not for someone generating shots on a laptop.

It runs in the browser. You describe a scene in plain language, for example "a dragon flies low over a village, banks around the tower and lands in the square while villagers scatter", and it lays out a 3D blockout: the set in simple shapes, the moving things animated, and a shot list with camera moves already in place. Then you adjust it by hand:

The studio is free: blocking, cameras, timing and export cost nothing. You pay only for the AI step that turns a sentence into a scene. New accounts get free scenes, and your first paid pack is $1 (₹9 in India) for three more.

The takeaway

If a video comes back with the wrong camera move, rewording the prompt a seventh time usually will not fix it. Words cannot carry that much information. Show the model the move instead of describing it: block it out, direct the camera, export the reference, and let the model do what it is good at.

Blockshot turns a plain-language description into a blockout with real camera moves. Try it free →