A model ships on a Tuesday. By Thursday someone on the team wants to try it, and the answer is a new account, a new API key, a new SDK, a new invoice and a new minimum balance. Do that four times and your "AI integration" is four integrations, each with its own auth scheme and its own failure mode — and nobody can tell you what last month actually cost, because the number is spread across four dashboards.
The Qtum AI Router exists so that the model name is the only thing that changes.
The whole integration is two environment variables
If your code already talks to OpenAI, you are done in one step:
export OPENAI_BASE_URL=https://router.qtum.ai/v1
export OPENAI_API_KEY=$QTUM_API_KEY
Any OpenAI-compatible SDK, CLI or tool now works unchanged. No wrapper, no adapter layer, no vendor-specific branch in your code. The raw call looks exactly like the one you already know:
curl https://router.qtum.ai/v1/chat/completions \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-sol",
"messages": [
{ "role": "user", "content": "Explain Qtum AAL in one paragraph." }
]
}'
Want a different model? Change the model string. That is the entire migration
path, and it is the reason this thing is worth having: the decision about which
model to use stops being an architectural commitment and becomes a config value
you can change on a Friday afternoon.
Both protocol families, one key
The Router speaks the Anthropic Messages API as well, on the same key and the same host:
curl https://router.qtum.ai/v1/messages \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-5-ab",
"max_tokens": 1024,
"messages": [
{ "role": "user", "content": "Explain Qtum AAL in one paragraph." }
]
}'
Which means the tools built on that protocol work too. For Claude Code and claude-code-router:
export ANTHROPIC_BASE_URL=https://router.qtum.ai
export ANTHROPIC_AUTH_TOKEN=$QTUM_API_KEY
One balance now pays for your agent sessions, your production inference and the model your colleague wanted to try on Thursday.
What is actually behind the endpoint
91 models at the time of writing, including 60 text, 13 image, 12 video and 3 audio. Not a curated shortlist of three: the current catalog includes the GPT-5.x and GPT-6 families, Claude Opus, Sonnet, Haiku and Fable, DeepSeek v4, Qwen 3.6–3.8, GLM 5.x, Kimi K2/K3, MiniMax, Seedream and Seedance, Wan, Sora 2 and HappyHorse.
That list changes, which is why this post is not the place to look it up. The live model table is generated from the catalog the gateway is actually serving, with the current price next to every model, and it is correct the moment a model is added or retired.
Pricing is per model — and that is the point
Prices are quoted in credits, where 1,000 credits = 1 USD. A few real figures from today's catalog, per million tokens:
| Model | Input | Output |
|---|---|---|
deepseek-v4-pro | 540 credits | 1,080 credits |
gpt-6-sol | 1,680 credits | 8,400 credits |
gpt-5.5 | 4,200 credits | 25,200 credits |
claude-opus-5-bedrock | 5,400 credits | 27,000 credits |
Look at the output column: twenty-five times between the cheapest and the most expensive row. That spread is the whole argument for a router. Classification, extraction and summarisation do not need your most expensive model; the hard reasoning step does. When switching costs you one string, you can actually act on that instead of paying frontier prices for everything because migrating is a sprint.
Concretely: a 2,000-token prompt with a 500-token answer on gpt-6-sol costs
about 7.6 credits — under one US cent.
Billing is prepaid and settles on real usage. A text request reserves its worst-case cost up front, then charges what the response actually consumed and returns the difference. There is no monthly commitment and no per-day request quota; the only gate is whether your balance covers the call.
The same key does images, speech and video
# Image
curl -X POST https://router.qtum.ai/v1/images/generations \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-image-2","prompt":"a clean product render of a red apple on a white table","size":"1024x1024","n":1}'
# Text to speech
curl -X POST https://router.qtum.ai/v1/audio/speech \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen3-tts-flash","input":"Hello from the Qtum router.","voice":"Cherry","response_format":"mp3"}' \
--output speech.mp3
# Video (async: create a task, then poll it)
curl -X POST https://router.qtum.ai/v1/videos \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-job-001" \
-d '{"model":"doubao-seedance-2-0-260128","prompt":"a calm ocean at sunset, cinematic","resolution":"720p","seconds":5,"mode":"text-to-video"}'
Same host, same key, same balance. The integration guide has the Python and TypeScript versions of each of these, plus the polling loop for video and the reference-file upload for image-to-video.
If you would rather not write any code
Qtum Video is the same video models behind a prompt box in the browser — no key, no SDK, no polling loop. Ten models, text-to-video and image-to-video, up to 30 seconds depending on the model.
Video is priced per five seconds of output, and the arithmetic is public:
ceil(credits_per_5s × seconds ÷ 5). So at 720p:
| Model | Per 5s | A 10-second clip |
|---|---|---|
doubao-seedance-1-5-pro-251215 (480p) | 80 credits | 160 credits ≈ $0.16 |
wan3.0-video | 347 credits | 694 credits ≈ $0.69 |
doubao-seedance-2-0-260128 | 994 credits | 1,988 credits ≈ $1.99 |
Two things worth knowing before you spend anything. A failed generation is refunded in full — credits are reserved when the task is created and returned if it does not produce a video, so a provider error costs you time and not money. And there is no daily or monthly generation limit; the only check is your balance. The full price table lists every model, every resolution and every clip length.
What it costs to find out
Nothing, to look: the model catalogue and the video prices are public pages, no account needed. When you want to make a call, the Router page is where you get a key — then top up and point your existing code at the base URL above. The first request should take about as long as reading this paragraph.
And if you only remember one thing: your model choice should be a config value. Everything above is in service of that.
