Video4 min read

What an AI video actually costs: ten models, one formula

AI video pricing is usually a pack of credits and a shrug. Here it is one multiplication you can do before you spend anything - and the same ten seconds ranges from 34 cents to three dollars depending on which of the ten models you pick.

Most AI video pricing is a guessing game. You buy a pack of credits, the pack buys an unstated number of "generations", and what a generation costs depends on a model, a resolution and a length that the pricing page does not multiply out for you. You find out what a clip costs by making one.

Qtum Video prices it as arithmetic you can do before you spend anything:

credits = ceil(credits_per_5s × seconds ÷ 5)

1,000 credits is 1 USD. Every model's credits_per_5s is listed per resolution on the price table, so the cost of any clip is a line of multiplication — no tiers, no packs, no minimum.

The same ten seconds, across every model

Ten seconds at 720p, which is the clip most people actually want:

ModelCreditsUSD
doubao-seedance-1-5-pro-251215344$0.34
doubao-seedance-1-0-pro-250528648$0.65
wan3.0-video694$0.69
happyhorse-1.1-t2v / -i2v892$0.89
happyhorse-1.01,600$1.60
doubao-seedance-2-0-2601281,988$1.99
doubao-seedance-2-5-2606283,026$3.03

Nearly nine times between the top and the bottom of that list, for the same length at the same resolution. Which is the real argument for having ten models instead of one: a storyboard pass and a final render are not the same job, and paying the final-render price to find out whether a shot works is how budgets disappear.

The cheapest thing you can make is a 5-second 480p clip on doubao-seedance-1-5-pro-251215 at 80 credits — eight cents. The most expensive 5 seconds is doubao-seedance-2-5-260628 at 1080p, 3,742 credits. Rough out the idea with the first; render it with the second.

Choosing a model

Price is one axis. Three others decide it for you.

How long the clip can be. doubao-seedance-2-5-260628 and wan3.0-video go to 30 seconds. Most of the rest stop at 10 or 15. sora-2-azure accepts exactly 4, 8 or 12 seconds; MiniMax-H3 takes any whole number from 4 to 15, which is the finest control in the catalog.

Whether you are starting from an image. Seven of the ten accept a starting frame. The four Seedance models go further and accept a last frame as well, so you can specify where a shot begins and ends and let the model fill in the motion between them. happyhorse-1.1 is split into two ids — -t2v for text and -i2v for image — at the same price, so pick the one matching your input.

Resolution. Most models offer 480p, 720p and 1080p. sora-2-azure is 720p only. MiniMax-H3 is the odd one out with 768p and 2K rather than the usual ladder.

Feeding it a starting image

On the models that accept one: JPEG, PNG and WebP everywhere, plus BMP and GIF on the Seedance models. Up to 30 MB on Seedance, 20 MB elsewhere. Each side at least 300px — 240px on wan3.0-video — and the aspect ratio has to sit between 0.4:1 and 2.5:1 on most models, where wan3.0-video is far more permissive at 0.125:1 to 8:1. On every one of these models a prompt is optional: hand it an image and it will animate it.

Every one of these limits is listed per model on the price table, read live from the same catalog the generator uses, so it cannot drift from what the API will accept.

A failed render is free

Credits are reserved when the task is created and returned in full if it does not produce a video. A provider error, a model that chokes on your prompt, a timeout — none of them cost you anything but the wait. That matters more than it sounds: it is what makes it reasonable to try the expensive model once rather than talking yourself out of it.

There is also no daily or monthly generation quota. The only check is whether your balance covers the clip. Ten short tests in an hour is a normal thing to do here.

Two ways in

In the browser. Qtum Video is a prompt box: pick a model, a resolution and a length, optionally drop in a starting image, and the cost is shown before you press the button. No SDK, no polling, nothing to install.

From code. The same models are on the API, through the same key as everything else on the Router. Video is asynchronous — create a task, poll it, fetch the result:

curl -X POST https://router.qtum.ai/v1/videos \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: my-first-clip" \
  -d '{
    "model": "doubao-seedance-2-0-260128",
    "prompt": "a calm ocean at sunset, cinematic",
    "resolution": "720p",
    "seconds": 5,
    "mode": "text-to-video"
  }'

That returns a task id with "status": "queued". Poll GET /v1/videos/{id} until it reports completed, then fetch the file from GET /v1/videos/{id}/content. The Idempotency-Key header means a retried create does not bill you twice — use a fresh one per distinct job. The integration guide has the Python version with the polling loop, and the reference-file upload for image-to-video.

Where to start

Open the price table, find the cheapest model that does what you need, and multiply. If the number is eight cents, there is not much to think about — go and make something.

One API key for every model

The Qtum AI Router puts text, image, audio and video models behind a single OpenAI- and Anthropic-compatible endpoint, billed in one place.