Text-to-video crossed the line from party trick to production tool somewhere in the last eighteen months. The clips are sharp, the physics mostly behave, and a ten-second shot that used to need a camera crew now needs a sentence.
Anyone who has actually used these tools knows the two places a project goes sideways.
Model choice. Every model has a personality. One holds a fluid camera move and melts faces by second twelve; another is cheap and fast and tops out at 720p. Render on the wrong one and you either pay flagship prices for draft-quality output, or wonder why your quick test came back looking like a trailer you did not need.
The prompt. "A dog running on a beach" is not a shot. A shot has a lens, a light source, a camera move and a mood — and video models respond to that vocabulary far more reliably than most people expect.
This guide covers both, with real prompts and the clips they rendered, inline, so you can judge for yourself. For what each model costs, see what an AI video actually costs; this one is about getting a good frame out the other end.
Pick the model before you write the prompt
Qtum Video currently runs ten generative video models behind one prompt box. You do not need to know all ten. You need to know which tier you are in:
- Drafting. Seedance 1.5 at 480p is the cheapest real render on the platform — 80 credits for five seconds, about eight cents. This is where you iterate.
- Everyday production. Seedance 1.5 at 720p or 1080p. Noticeably better prompt adherence than 1.0 — it actually respects your lens and lighting direction — with stable subjects across the clip.
- Final renders. Seedance 2.0 and 2.5. Cinematic detail, convincing physics, orbits and dolly-ins that hold together, and multi-shot continuity. Seedance 2.5 at 1080p costs roughly forty-seven times a 480p draft, which is the entire argument for drafting cheap and finishing expensive.
- Long clips on a budget. Wan 3.0 and Seedance 2.5 both reach thirty seconds. HappyHorse 1.1 and MiniMax H3 reach fifteen for considerably less than the Seedance flagship.
- Starting from a still. Seven of the ten accept a first frame; the four Seedance models accept a last frame as well, which is how you pin both ends of a move.
| Model | Max clip | Resolutions | Starting frame |
|---|---|---|---|
| Seedance 2.5 | 30 s | 480p · 720p · 1080p | first and last |
| Seedance 2.0 | 15 s | 480p · 720p · 1080p | first and last |
| Seedance 1.5 | 10 s | 480p · 720p · 1080p | first and last |
| Seedance 1.0 | 10 s | 480p · 720p · 1080p | first and last |
| Wan 3.0 | 30 s | 480p · 720p · 1080p | first |
| HappyHorse 1.0 | 15 s | 720p · 1080p | first |
| HappyHorse 1.1 I2V | 15 s | 480p · 720p · 1080p | required |
| HappyHorse 1.1 T2V | 15 s | 480p · 720p · 1080p | — |
| MiniMax H3 | 15 s | 768p · 2K | — |
| Sora 2 | 12 s | 720p | — |
Rule of thumb: draft on Seedance 1.5 at 480p, produce at 720p, finish on 2.0 or 2.5 — and reach for Wan 3.0 or HappyHorse when the clip is long or the budget is tight. The models and pricing page always shows the live figures.
The anatomy of a shot prompt
Video models are trained on footage that came with descriptions written by people who think in shots. The closer your prompt reads to a shot description from a treatment or a stock-footage caption, the better the model performs. The single biggest upgrade available to you is to stop describing a thing and start describing a shot.
A reliable order:
Subject → action → setting → camera and lens → lighting → style and mood
A weathered fisherman in a yellow oilskin hauls a net over the gunwale, spray flying, on a small boat in rough grey-green seas. Handheld 35mm, tracking close on his hands, overcast diffuse light, cold and even. Gritty documentary realism, muted colour grade.
Front-load what matters: models weight the start of the prompt heaviest. Subject and action first, garnish later. And keep it to one scene per prompt — if your prompt contains "and then", you are describing an edit, not a shot. (The exception is the multi-shot models, which understand explicit cut to: transitions. There is an example below.)
Speak lens. Speak light.
These two vocabularies do more work per word than anything else you can type. You do not need film school; you need about a dozen terms.
Camera and lens
| Term | What it buys you |
|---|---|
wide-angle / 16mm | Big environments, dramatic perspective, landscapes and interiors |
35mm | The documentary look — natural perspective, what your eye expects |
85mm, shallow depth of field | Portrait compression, creamy blurred background, subject isolation |
macro | Extreme close-up: texture, droplets, mechanical detail |
aerial drone shot | High and moving. Pair with "slowly orbiting" or "flying over" |
dolly-in / dolly-out | Smooth push toward or pull away from the subject |
tracking shot | Camera travels alongside a moving subject |
handheld | Subtle shake, documentary energy. Omit it and you get tripod-smooth |
static shot | Underrated: it stops the model inventing camera moves you did not ask for |
Lighting
| Term | What it buys you |
|---|---|
golden hour | Warm low-angle sun, long shadows. The most flattering light there is |
blue hour | Just after sunset — cool, moody, city lights starting to glow |
overcast / diffuse | Soft, even, shadowless. Good for faces and product shots |
rim lighting / backlit | Bright edge around the subject, separating it from the background |
volumetric light / god rays | Visible beams through mist, dust or windows. Instant atmosphere |
neon / practical lights | Light from sources inside the scene. Pairs beautifully with rain |
candlelight / tungsten | Warm, orange, intimate, flickering |
high-key / low-key | Bright and airy versus dark and dramatic — the emotional register in one word |
Six habits that consistently pay off
- Concrete nouns beat adjectives. "A rusted blue 1970s pickup" renders better than "an old cool truck". Models cannot paint "cool"; they can paint rust.
- Give every subject a motion verb. Video models need to know what moves. Steam rising, hair blowing, waves crashing. Static prompts produce static, lifeless clips.
- Describe what you want, never what you don't. Negations backfire: "no people in the street" often renders people. Say "an empty street".
- Name one style anchor, not five. "Shot on 35mm film" or "anamorphic lens flare, cinematic" sets a coherent look. Stacking "cinematic, anime, hyperrealistic, vaporwave" gives the model a contradiction, not a style.
- Specify pace. Slow motion, time-lapse, real-time — otherwise the model picks for you.
- Keep it to 40–90 words. Long enough to direct, short enough that nothing gets ignored.
A fifteen-second makeover
Before — describes a thing:
A cool dog running on the beach at sunset, very beautiful, high quality, 4K, epic.
After — describes a shot:
A golden retriever sprints along the wet sand at the water's edge, ears flying, kicking up spray. Low tracking shot moving alongside, 85mm with shallow depth of field, golden hour backlight, warm rim light on the fur. Joyful, cinematic, slow motion.
Same dog, same beach. The second tells the model where the camera is, what the light is doing and how time flows. Quality keywords like "4K" and "epic" do almost nothing; lens and light do almost everything.
Four prompts, rendered
Every clip below was rendered as-is on Qtum Video with default settings — no cherry-picking. The first shows a single model; the next three run the same prompt through two models side by side, so you can see exactly what the extra credits buy. Copy the prompts, remix them, re-render them yourself.
1 · The fast draft
A single subject, tight framing, a few seconds — exactly the kind of shot you iterate on cheaply before spending flagship credits.
Macro close-up of an espresso pouring into a glass cup, rich crema
swirling and folding. Warm café bokeh in the background, soft morning
window light, shallow depth of field. Slow motion, cozy and inviting.
2 · The everyday upgrade — 1.0 against 1.5
More demanding: a moving subject, a moving camera, practical lighting and a specific colour grade. The same prompt on both models shows what the mid-generation jump actually buys — tighter prompt adherence and steadier reflections.
A cyclist in a yellow raincoat rides through a rainy Tokyo street at
night, neon signs reflecting in the puddles. Tracking shot moving
alongside, 35mm lens, light rain falling through the glow of the signs.
Teal-and-magenta cinematic color grade, moody and atmospheric.
Watch the puddle reflections and the rain holding together under the tracking move on 1.5 — the kind of direction 1.0 starts to lose. For a little more per clip, 1.5 is the sensible everyday default.
3 · The money shot — 1.0 against 2.0
Everything that breaks lesser models, in one prompt: water physics, an orbiting aerial camera, volumetric light and a film look. Run on the entry model and the flagship back to back, the gap is the clearest argument for when flagship credits are worth spending.
Aerial drone shot slowly orbiting a white lighthouse on a rugged cliff
at golden hour. Waves crash against the rocks below, sea mist drifting,
volumetric god rays breaking through the spray. Anamorphic lens flare,
epic cinematic mood, 24fps film look, rich warm color grade.
The crashing water and the continuous orbit are exactly what separates a flagship from the rest. 1.0 gets you the idea; 2.0 holds the physics and the camera move together.
4 · The long multi-shot
A fifteen-second, three-shot micro-story using explicit cut to: transitions. Both models handle multi-shot continuity, so the comparison is about price and personality.
A street-food vendor flips a sizzling pancake on a griddle, steam
rising into warm tungsten market lights. Cut to: hands drizzling sauce
and folding the pancake into paper. Cut to: the pancake handed across
the counter to a waiting customer, market bustle blurred behind.
Handheld 35mm documentary style, shallow depth of field, warm and lively.
Three shots, fifteen seconds, one render each. HappyHorse holds the motion and the cuts at a meaningfully lower cost; Seedance 2.0 keeps faces and fine textures crisper across the shots. Pick by what the clip needs — length and budget, or detail.
The cheapest way to learn a model's personality is to render the same prompt on two of them and watch where they diverge. Do it at the lowest resolution and the shortest length: a five-second 480p draft costs a small fraction of a 1080p one, and the personality difference shows up just as clearly.
What it costs, and what happens when a render fails
Qtum Video skips the subscription playbook. No monthly plan, no seat licence, no credit packs that expire at the end of the quarter.
- Credits are the unit of work. Every render quotes its cost up front — from the model, the clip length and the resolution — and deducts exactly that.
- QTUM is the primary way to buy credits. Top up straight from your wallet; MetaMask Snap makes it a few clicks. No card on file, no fiat checkout, no billing relationship to manage.
- A failed render is refunded. Credits are reserved when a render starts and returned in full, automatically, if it does not complete. You pay for finished footage, not for attempts.
Pay-as-you-go changes how you work. There is no monthly quota to use up, so there is no pressure to over-render; and because a flagship render visibly costs more than a draft, the draft-cheap / finish-expensive workflow in this guide is not just good practice, it is the economically obvious one. The cost breakdown has the per-model arithmetic.
No credit card also means no card-linked identity. Sign in with MetaMask Snap, pay in QTUM, and the service knows your wallet and your prompts and nothing else. Your renders are never used to train the models.
Pick a prompt off this page — they are written to be stolen — choose a model, and watch it come back as footage.
Open the studio → · Models and pricing → · Video over the API →
Every example clip was rendered on Qtum Video from the prompt shown, at default settings. Hero photograph by Taylor Vick on Unsplash.
