Fellowi

Text to Video or Image to Video: Which One You Actually Want

By The Fellowi Team · · 6 min read

Abstract artwork of two parallel light paths converging on a bright frame against a dark field

Fellowi runs four video models. Two of them start from a sentence and two of them start from a picture, and each of those pairs comes in a fast turbo version and a slower full one. That is the whole catalogue, and the only choice that really changes what you get back is the first one: are you handing the model a still, or a description?

We should admit that we described this badly for a while. Three posts on this blog stated flatly that there was no text-only path, which was wrong, and we corrected them. If you read something here in early September saying image to video is the only option, that is the sentence we fixed.

The one rule that settles most cases

If a specific face has to come out the other end, start from an image.

This is not a preference. A text description cannot carry a face. “A woman in her late twenties with dark hair and green eyes” narrows the category a little and then leaves thousands of faces inside it, and the model picks a different one each run. Handing it a still removes the guessing entirely: the person in frame is already decided, and the model's job shrinks to moving them. We wrote about why identity lives in a reference image rather than in words, and video is the sharpest version of that problem, because inconsistency across twenty-four frames a second is far more visible than across two separate pictures.

Text to video earns its place when the subject is generic and the scene is the point. Waves hitting a breakwater, traffic at night, a candle burning down, an establishing shot of a street. Nobody in the audience knows what those are supposed to look like, so there is nothing to be inconsistent with, and you skip the entire step of making a still first.

What each path actually asks of you

Text to video wants a prompt that carries the whole world: subject, place, light, camera, and motion. It is the harder prompt to write because nothing is pinned down for you, and a vague one gets you an averaged, stock-looking clip.

Image to video wants almost the opposite. The still has already said what the scene looks like, so repeating any of that in the prompt is wasted, and contradicting it is worse than wasted. Spend every word on motion instead: what moves, how fast, and what the camera does. Our guide to video prompts is built around exactly that discipline, and one of our paste-ready image prompts is written specifically to be a good start frame.

There is a third thing the image path can do that the text path cannot: take an end frame as well as a start frame. Give it both and it renders the motion between two stills you picked, which is the most control over a clip available here. It is the natural tool for a short vertical where you know exactly where the shot has to land.

Turbo or full, which is a separate question

Once you have picked a path, you still pick a family, and this one is about ceiling and price rather than fidelity of concept.

The turbo models are the fast, cheap pair. They offer 720p and 1080p, and a four-second 720p clip is 500 coins. The full models are slower and go further: they add 480p at the bottom and 4K at the top, and their 720p is visibly better, which is why it costs 900 coins for the same four seconds rather than 500. The gap widens fast with resolution, because these are per-second prices and 4K is an enormous amount of pixels: four seconds of full 4K is 4,400 coins, and thirty seconds of it is 33,000.

The honest advice is to draft in turbo. Get the motion and the framing right at 720p for 500 coins, and only then decide whether this particular clip deserves a full-model render. Every option is priced per second, so the cost of finding out you wanted a different camera move is entirely up to which model you found out on.

The settings both paths share

Duration runs from four to thirty seconds. There are seven aspect ratios, including 9:16 for vertical social video, plus an auto setting that keeps the start image's own framing rather than imposing a shape on it. Audio is a checkbox, generated along with the clip, and it does not change the price. There is more on every dial in the post about the control surface.

One practical note on price: the cost is linear in duration and set by model and resolution, so it is identical whichever path you came in by. Choosing text to video does not save you anything, and choosing image to video does not cost extra. Pick on the merits.

A generation that fails returns your coins automatically. So the cheap experiment is the one worth running first: make a still in Fellowi Images for 40 coins, take it to Fellowi Video, and describe only what moves. If the subject turns out not to need a face at all, you have just learned you could have skipped a step, which is a cheap thing to learn.

Try it for yourself

A warm, private AI companion - 7 days free with 30 messages, no card needed.

Pricing and limits

Keep reading