How to Write AI Video Prompts: Describe the Motion, Not the Picture
By The Fellowi Team · · 7 min read

Here is the one idea that fixes most bad video prompts: in image-to-video, the image already did half the work. Whatever you upload to Fellowi Video settles who is in the frame, what they wear, where they stand, how the light falls. The model can see all of it. So when your prompt spends thirty words repeating that, you have described the starting point twice and the next four seconds not at all. The image does the what. The prompt does the how it moves.
If you have read our guide to writing image prompts, notice how different the job is. An image prompt builds a scene from nothing, so detail is your friend. A video prompt receives a finished scene and adds time. Detail about the scene is now noise; detail about motion is the signal.
The five things a video prompt should answer
You do not need all five in every prompt. But every good prompt answers at least two or three, and a weak one usually answers none.
1. Camera: does it move, and how?
Static, slow push-in, slow pull-back, gentle pan left or right, a soft handheld drift. Name one. A camera that moves slowly is the easiest way to make a still feel like footage, even if nothing else changes.
2. Subject motion: what does the person or object do?
A turn of the head, a breath, a smile that arrives, hair lifting in a breeze, fingers closing around a cup. Small and specific beats big and vague. “She looks up and smiles” will land; “she moves naturally” leaves the model to guess.
3. Pace: how fast is all this happening?
Slow motion, real time, a lazy drift, a quick glance. Pace decides whether four seconds feel like a moment or a montage. When in doubt, say “slow”. Slow reads as intentional; fast reads as chaos.
4. Mood and environment: what changes in the world?
Light flickering, curtains stirring, rain starting, steam rising from coffee, dust in a sunbeam. One environmental detail gives the scene a pulse and stops the background from looking pasted on.
5. Sound: what should we hear?
Fellowi Video generates audio with every clip, and the prompt is where you steer it. Soft rain, distant traffic, a quiet room with a clock, wind through leaves, low ambient music. If you say nothing, the model picks. If you say one thing, it usually listens.
Before and after
The same start image in each pair. Only the prompt changes.
A portrait by a window
Before:“A beautiful woman with long dark hair in a white shirt standing by a window in a bright apartment.”
That is a caption of the image. The model already knows every word of it, and nothing here says what happens.
After:“Slow push-in. She turns her head towards the window and exhales softly; a strand of hair lifts in a light breeze. Curtain sways. Quiet room, faint city hum outside.”
Camera, subject motion, environment and sound, in one breath, with nothing repeated from the picture.
A mountain lake at dusk
Before:“Epic cinematic sunset over a lake with mountains, dramatic lighting, 4K, hyperrealistic.”
Adjectives about quality do nothing for motion, and “4K” is not a thing you can ask a 720p clip to be.
After:“Static camera. Ripples spread slowly across the water; clouds drift right; the last light on the peaks dims a little. Wind over water, a single distant bird.”
A generated photo from the Photo tab
Before:“Make it move. Realistic. She walks across the room, picks up her phone, sits on the sofa, laughs and waves at the camera.”
Four actions in four seconds. The model will attempt all of them and finish none.
After:“Handheld drift, slight. She glances at the camera and breaks into a slow smile. Soft indoor light flickers. Muffled music from another room.”
One action with a beginning and an end, and it fits comfortably in the clip. If you also want the walk to the sofa, that is a second clip with a different start image, not a longer sentence. Our post on animating your AI photos walks through building a sequence this way.
Three mistakes that ruin a clip
Describing the image again
The most common one, and the most forgivable, because image prompting trains you to do exactly this. Test yourself: cover the picture and read your prompt. If it could be a caption for the still, you have not written a video prompt yet. Cut every word the model can already see and spend them on motion instead.
Too many actions for four seconds
A head turn takes about a second. A smile takes another. A step or two takes the rest. That is the whole clip. Anything past two clear movements gets compressed into a blur or silently dropped. Pick the one gesture the moment is about and let the camera and the environment carry the rest.
Contradicting the image
If the person is sitting, do not ask them to run. If it is night, do not ask for bright sunlight. If they face away, do not ask for a close-up of their eyes. The model has to honour the start frame, so a prompt that fights it produces warping, morphing, or a face that stops looking like the face you uploaded. Work with what is in the picture: the motion you ask for should be one that this exact frame could plausibly lead into.
A small template to start from
Fill in what applies and delete the rest:
- Camera: static / slow push-in / slow pull-back / gentle pan / handheld drift
- Subject: one gesture with a start and an end
- Pace: slow / real time
- Environment: one thing that moves in the background
- Sound: one or two ambient cues
Written out: “Slow pull-back. He lowers the book and looks out at the rain. Lamplight flickers. Rain on glass, pages settling.” That is about twenty words, and it is a complete video prompt.
Frequently asked questions
How long should a video prompt be?
One to three short sentences. If you are past forty words, you are probably describing the picture or stacking actions. Cut until only motion, pace and sound remain.
Should I mention what the person looks like?
Only if it is part of the motion (“her hair lifts”, “his scarf flutters”). Their appearance is fixed by the image; restating it changes nothing and crowds out the useful words.
What if the render fails?
A failed render is never charged: coins are refunded and a Director plan slot is returned. The usual fix is to simplify: fewer actions, a slower pace, and no instruction that argues with the start image. Try again with the same picture.
Can I ask for explicit content?
Fellowi is an 18+ platform, and adult content between adults is allowed. Anything involving minors or non-consent is refused outright. The prompting rules above apply exactly the same way: describe the motion, keep it to one or two beats, and do not contradict the frame.
Where do I try this?
Open the studio from the magic-wand icon in the header, or go straight to fellowi.com/video. Pick a start image, paste your prompt, and watch what twenty well-chosen words can do.