Fellowi

Making One Short Clip, From Blank Page to Finished File

By The Fellowi Team · · 8 min read

Abstract artwork of a single band of blue light stretching and blurring across a dark frame

I wanted one ten-second clip: a woman on a balcony at night, city behind her, the kind of shot that would open something. Nothing ambitious. It took about forty minutes and 3,150 coins, and writing down where those went is more useful than any tidy tutorial, because most of the forty minutes was not spent typing.

Deciding the path first, which I nearly got wrong

My instinct was to describe the whole thing in text and let the model invent it. That would have been the wrong call, and the reason is the one thing worth knowing before you start: text to video invents the person too, and a different one every run.

Since I wanted a specific look that would survive the whole clip, the answer was to make a still first and animate that. If I had wanted waves on a breakwater or traffic at night, text would have been fine and I would have saved a step. That whole decision is one sentence long and the comparison post settles it.

The still: four images, 160 coins

Standard quality, 40 coins each. First was too bright, second had her looking at camera when I wanted her looking away, third was right but framed too wide, fourth was the one.

One thing I did deliberately, having learned it the expensive way before: I wrote the still to be a start frame, not a good photograph. That means the subject is at rest rather than mid-gesture, because a still that already looks like the middle of a movement gives the video model nowhere to go. And the background is out of focus, because backgrounds are where drift shows up first.

The first clip: four seconds, 500 coins, wrong

I drafted at four seconds and 720p on a turbo model, which is the cheapest useful combination there is. My prompt described the scene: the balcony, the city lights, her hair, the mood.

It came back looking like the still had been mildly disturbed. Barely any motion, and what motion there was felt aimless.

The mistake is obvious in retrospect and I make it every time I have been away from video for a week. The still had already said what the scene looks like. Repeating that in the prompt spent my entire instruction budget on information the model already had. A video prompt should be about what moves.

The second clip: four seconds, 500 coins, nearly

Rewrote it to describe motion only: a slow breath, her hair lifting slightly in the wind, the city lights softly shifting focus behind her, camera static. Four clauses, all of them about change over time.

Enormous difference. Now it was a shot rather than a wobbling photograph. Still not right, though: the camera being static made it feel like security footage. Which was useful to learn, because it meant the camera was the variable, and I had not touched it.

The third clip: four seconds, 500 coins, that is the one

Changed exactly one thing. Static camera became “very slow push in”. Everything else identical.

That was the shot. A slight movement toward the subject does something to the attention that nothing in the frame can do by itself, and it cost one clause.

Same discipline as images: one variable per attempt. It is more important with video, not less, because each attempt costs twelve times as much.

Committing to length: ten seconds, 1,250 coins

Only now, with the motion and the camera settled, did I re-run at ten seconds. Same prompt, same start frame, longer duration.

This is the part of the workflow that saves the actual money. Ten seconds is 1,250 coins at 720p turbo; I ran the experiments at 500 and paid the ten-second price once, at the end. Getting that backwards is how a single clip eats a large pack, and the pricing ladder is steep enough that the discipline pays for itself immediately.

One honest note: the longer version is not simply the four-second one extended. It re-generates, so the motion is similar but not identical. Occasionally the longer take is worse and you run it again. Budget for that possibility rather than being surprised by it.

Sound: nothing

Literally nothing to do. The clip arrived as one MP4 with an audio track already inside it, generated for the same scene, and the price did not change. I had written “distant traffic, wind” into the prompt and that is roughly what was there.

This is the step that used to be half the work, and it is now the step that does not exist. More on why in the post about generated audio.

The final tally

  • 4 stills at 40 coins: 160
  • 3 draft clips at four seconds: 1,500
  • 1 final clip at ten seconds: 1,250
  • Total: 3,150 coins, about forty minutes

Of those forty minutes, maybe six were typing. The rest was waiting for renders, which run for minutes rather than seconds, and standing around deciding whether the last one was good enough. That rhythm is the real difference from image work: you cannot iterate rapidly, so each attempt has to be worth making.

What I would tell myself before starting: write the still as a start frame, describe only motion, change one variable, and do not touch duration until the four-second version is the shot you want. Everything else is patience. Fellowi Video is where it happens, and the video prompt guide is the compressed version of the two clips I wasted.

Try it for yourself

A warm, private AI companion - 7 days free with 30 messages, no card needed.

Pricing and limits

Keep reading