← Blog

2026-09-11

Motion Prompts for AI Video: Subject, Camera, Rhythm

Motion Prompts for AI Video: Subject, Camera, Rhythm

A good motion prompt for AI video describes three things separately: what your subject does, what the camera does, and the pace at which it happens. Blend those layers together and the model has to guess what should move, and that guessing is exactly what causes warped faces and melting backgrounds. Here is how to build a motion prompt that stays stable.

What is a motion prompt, exactly?

A motion prompt is the part of your prompt that describes what moves, separate from the still scene. With image-to-video your frame is already fixed: subject, clothing, light and composition all come from the start image. Your prompt does not need to restate those, only to say how the frame comes to life.

  • For image-to-video: spend your words on the movement, not on the scene.
  • For text-to-video: set the scene briefly, then use most of your words on motion.

The more clearly you separate "what is visible" from "what moves," the less the model has to guess.

The three layers: subject, camera and environment

Split every motion prompt into three separate layers, so the model knows what moves independently.

  • Subject motion: what your character or product does. "She slowly turns her head toward the camera and smiles."
  • Camera motion: how the camera moves, in film terms the model recognizes: dolly in, pan, tilt, tracking shot or a static locked-off shot.
  • Environmental motion: everything around it that is alive, such as wind in the hair, waves, drifting dust, rising steam or shifting light.

Name the subject and its action first. Modern video models weigh the start of your prompt most heavily, so your most important movement belongs up front.

Why one main movement per clip works best

Keep it to one clear main movement per clip. Most models render short fragments: Kling generates 5 or 10 seconds per generation, with 5 seconds as the default. A string of separate actions simply does not fit into those few seconds.

Stack too much movement and the frame starts to morph. Many video models do not build a true 3D model of the scene; they learn from flat clips and cannot always tell a panning camera apart from an environment that changes shape. Combine a fast whip pan with a macro zoom, for example, and you often get a blurry mess. One camera move plus one action keeps your subject recognizable.

Describe it concretely, not vaguely

Write physical actions, not mood. Words like "dynamic," "epic" or "cinematic" tell the model nothing about what should move; a concrete description does.

  • Vague: "a dynamic, cinematic shot of a runner."
  • Concrete: "a runner sprints from left to right across the frame, the camera follows him in a tracking shot, dust kicks up behind his shoes."

Describe the speed too. "Slowly," "gradually" or "in one smooth motion" stops the model from overdoing the movement. Not sure how to phrase it? The prompt generator turns your idea into a structured prompt with scene, camera and motion built in.

Chaining several movements cleanly

If you do want more than one movement in a single clip, use timing words as sequence cues. Words like "starts," "then" and "as" act as edit points inside a single generation.

Example: "The camera holds still on her hands, then slowly tilts up to her face and dollies in as she starts to speak." Want a longer scene? Render short clips separately and place them back to back, or let the video continue from the last frame using video extend. That way you keep control per fragment and every movement stays clean.

Frequently asked questions

Should I describe the scene again for image-to-video?

No. With image-to-video your frame is already fixed in the start image. Describe only the movement; restating the scene distracts the model and raises the chance of distortion.

How many movements can I put in one prompt?

Keep it to one main movement per clip. More rarely fits believably into the five to ten seconds most models render per generation, and it raises the chance your subject starts to morph.

Why does my subject distort during a camera move?

Many models lack a true sense of 3D space and confuse a moving camera with an environment that changes shape. One clear camera move and one concrete subject action keep the frame stable.

Do vague words like "cinematic" work?

Barely. The model cannot attach any movement to them. Instead, describe the concrete action, the camera movement and the pace.

A strong motion prompt is mostly a matter of separating and getting specific. Want to try it? Create an account and bring your first frame to life with the video generator.