"Video AI" covers a lot of ground: text-to-video tools that invent a scene from a paragraph, lip-sync tools, background removers, and the image-to-video models that animate a photo you already own. This site belongs to the last group, and the difference matters more than the marketing suggests. Here is what an image-to-video model actually does, what you control, and where the money goes.
The two kinds of video AI
Most consumer tools fall into one of two buckets. A text-to-video model takes a written prompt and generates a scene from nothing; nobody in the result has your face, and the likeness you get is whatever the model decided to draw. An image-to-video model takes a photo you supply and animates it, so the subject in the clip is the person or animal in your picture.
Both are useful, but they solve different problems. If you want a generic cinematic shot, text-to-video is the right shape. If the point is that the face is recognisable, you need image-to-video, because the identity has to come from somewhere, and a prompt cannot supply it.
What image-to-video changes
The input is a single photo. The model reads the face, the pose, and the lighting in it, then builds a short clip that keeps the subject recognisable while the scene plays out. On this site the clip is usually five seconds, framed for a phone screen.
That short length is not a limitation so much as a decision. Five seconds is enough for one emotional beat or one visual turn, and it is the length the short-video feeds are built around. The effects here lean into it: each one is a small story with a clear beginning and a payoff, not a long scene that needs a lot of setup.
What you actually control
With an image-to-video model, the levers are fewer than people expect, and they are mostly about the photo.
You control the face, because you choose which picture goes in. You control the light, because the source photo's lighting carries into the clip. You control the framing, because a face that fills the frame survives the animation better than a tiny figure in a wide landscape. And you control which effect runs, because each of the five shapes the story differently.
What you do not control is the fine detail of every frame. The model supplies the motion, the setting, and the transitions. That is why the practical skill here is photo selection rather than prompt writing. The effects on this site are image-to-video effects, so the best results come from treating the photo as the script.
The five effects in practice
The site runs five effects on one pipeline, and moving between them does not mean learning a second tool.
The zombie love story uses a photo of a person and plays a four-beat scene: the standoff, the lowered weapon, the hug, and the hard cut into a warm memory. The zombie pet effect does the same kind of turn with an animal. The AI Halloween makeup effect puts a full horror look on a face. The AI mini me effect shrinks the subject into the frame, and the make-things-giant effect scales a subject up against its surroundings.
Once you know how to pick a photo for one, you know how to pick a photo for all of them: clear face, even light, one subject, no heavy filters.
Where the cost sits
A clip here is 300 credits, and a clip is five seconds. The smallest pack is $4.99 and covers exactly one clip. $9.99 gives 1,500 credits for five clips, and $19.99 gives 3,300 credits for eleven. The same three prices are the monthly plans, and there are yearly options as well. The full set is on the pricing page.
Two things are worth stating plainly, because a lot of video AI marketing is vague about them. Video here is paid from the first clip, and there is no free trial for a clip — the one free run a new account gets covers an image effect, not a video. What you get instead is a price shown before you generate, so an attempt never costs more than you were told, and a downloaded file with no watermark on it, so the clip you post is the clean version.
What a clip is not
An image-to-video clip is not a deepfake of a stranger, and it is not a video editor. It does not let you swap one face onto another person's body from a stock clip, and it does not give you frame-level control. It animates the photo you bring. Treated as a fast way to make five seconds of cinematics out of a picture you already have, it is genuinely useful. Treated as a full production suite, it will disappoint.
The honest framing is this: you supply the face and the photo quality, the model supplies the scene. When both halves are good, the result reads as a real clip rather than an effect demo. When the photo is weak, no amount of prompt writing rescues it.
Try it with one photo
The cheapest way to understand image-to-video is to run one clip. Pick a clear photo, upload it to the zombie video generator, and check the credit cost before you generate. If you want ideas for what to build first, the zombie AI video prompts list fifteen copyable scene briefs, and the pricing page shows what a clip and a bundle cost.