PhotoLoop
AI Text to Video Generator
Describe a scene and get video back. Every model's real cost and render time is published below, before you spend a credit.
Nothing to anchor to, for better and worse
Generating from text removes the constraint that makes image-to-video predictable. There is no first frame to stay consistent with, so the model invents the subject, the setting and the light. That is the appeal, and it is also why two runs of the same prompt give you two different scenes. Plan to generate more than once.
Write for the four channels the model reads
These models were trained on video paired with detailed description, so a vague prompt returns a vague, averaged result. State the camera path, the subject action, the setting and the ambient sound as separate clauses. Our prompt optimiser applies the official per-model rules rather than a generic rewrite, and it declines to use Seedance-specific techniques on models whose makers never documented them.
What each model costs for five seconds
At 720p: Wan 3.0 is 38 credits, Kling 3.0 is 60, Seedance 2.0 is 100 and Seedance 2.5 is 148. At 1080p: 75, 80, 249 and 266. One credit equals one cent of supplier cost, so those figures are the compute, not a markup. Wan 3.0 is deliberately the default first step because it is the cheapest way to find out whether the idea works at all.
Expect two to four minutes
Video inference is slow everywhere, not just here. Our measured range is 100 to 220 seconds. You can leave an email and close the tab rather than watching a spinner, and the model list shows the estimate next to each option before you choose.
- Four models with published per-clip costs
- 5 or 10 seconds, 720p or 1080p
- Prompt optimiser follows each model's own rules
- Requires credits — the free daily clip is image-to-video
Questions
How is text to video different from animating a photo?
With a photo the model is constrained: it has a real first frame and mostly has to stay consistent with it. From text alone it invents everything, which means more freedom and far less predictability. The same prompt run twice gives you two different scenes.
What makes a prompt work?
Separate the four things the model handles independently: what the camera does, what the subject does, what the setting looks like, and what it should sound like. 'A quiet street at dusk' produces an average of every dusk street. 'Camera static, steam drifting from a vent, a neon sign flickering, one passer-by walking out of frame, soft city ambience' produces a shot.
Should I use timeline segments?
Only on Seedance models. Their documentation explicitly supports splitting a prompt into 0-3s and 3-6s beats, and it helps on longer or more complex motion. Wan and Kling have no such published guidance, so we do not apply that technique to them rather than guess.
Is there a free tier for text to video?
No. The free daily clip is image-to-video only, because the free model is locked to Seedance 2.0 Mini in its image-to-video form. Text to video needs credits: 38 for Wan 3.0 at 720p, up to 266 for Seedance 2.5 at 1080p for a 5-second clip.
Which model handles text prompts best?
Seedance models are the ones with published prompt-writing guidance, which is why our prompt optimiser applies their rules by default. Kling 3.0 generates audio. Wan 3.0 is the cheapest way to see whether an idea holds up at 38 credits for 5 seconds at 720p.
Can I get a longer clip?
Ten seconds is the other supported length, and it costs exactly double — 75 credits on Wan 3.0 at 720p instead of 38. Pricing is per second of output, so there is no discount for length.
Does it generate sound?
Kling 3.0 does, and our pricing assumes audio is on because that is the default here. If you turn audio off you are charged about a third more than necessary — a deliberate conservative rounding on our side rather than an attempt to undercharge and lose money.
Keep exploring
PHOTOLOOP