Router4 min read

One API key for every model: a practical tour of the Qtum AI Router

Point the OpenAI or Anthropic SDK you already use at one endpoint and reach 91 models - text, image, audio and video - on a single prepaid balance. The whole integration is two environment variables, with real prices.

A model ships on a Tuesday. By Thursday someone on the team wants to try it, and the answer is a new account, a new API key, a new SDK, a new invoice and a new minimum balance. Do that four times and your "AI integration" is four integrations, each with its own auth scheme and its own failure mode — and nobody can tell you what last month actually cost, because the number is spread across four dashboards.

The Qtum AI Router exists so that the model name is the only thing that changes.

The whole integration is two environment variables

If your code already talks to OpenAI, you are done in one step:

export OPENAI_BASE_URL=https://router.qtum.ai/v1
export OPENAI_API_KEY=$QTUM_API_KEY

Any OpenAI-compatible SDK, CLI or tool now works unchanged. No wrapper, no adapter layer, no vendor-specific branch in your code. The raw call looks exactly like the one you already know:

curl https://router.qtum.ai/v1/chat/completions \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-sol",
    "messages": [
      { "role": "user", "content": "Explain Qtum AAL in one paragraph." }
    ]
  }'

Want a different model? Change the model string. That is the entire migration path, and it is the reason this thing is worth having: the decision about which model to use stops being an architectural commitment and becomes a config value you can change on a Friday afternoon.

Both protocol families, one key

The Router speaks the Anthropic Messages API as well, on the same key and the same host:

curl https://router.qtum.ai/v1/messages \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-5-ab",
    "max_tokens": 1024,
    "messages": [
      { "role": "user", "content": "Explain Qtum AAL in one paragraph." }
    ]
  }'

Which means the tools built on that protocol work too. For Claude Code and claude-code-router:

export ANTHROPIC_BASE_URL=https://router.qtum.ai
export ANTHROPIC_AUTH_TOKEN=$QTUM_API_KEY

One balance now pays for your agent sessions, your production inference and the model your colleague wanted to try on Thursday.

What is actually behind the endpoint

91 models at the time of writing, including 60 text, 13 image, 12 video and 3 audio. Not a curated shortlist of three: the current catalog includes the GPT-5.x and GPT-6 families, Claude Opus, Sonnet, Haiku and Fable, DeepSeek v4, Qwen 3.6–3.8, GLM 5.x, Kimi K2/K3, MiniMax, Seedream and Seedance, Wan, Sora 2 and HappyHorse.

That list changes, which is why this post is not the place to look it up. The live model table is generated from the catalog the gateway is actually serving, with the current price next to every model, and it is correct the moment a model is added or retired.

Pricing is per model — and that is the point

Prices are quoted in credits, where 1,000 credits = 1 USD. A few real figures from today's catalog, per million tokens:

ModelInputOutput
deepseek-v4-pro540 credits1,080 credits
gpt-6-sol1,680 credits8,400 credits
gpt-5.54,200 credits25,200 credits
claude-opus-5-bedrock5,400 credits27,000 credits

Look at the output column: twenty-five times between the cheapest and the most expensive row. That spread is the whole argument for a router. Classification, extraction and summarisation do not need your most expensive model; the hard reasoning step does. When switching costs you one string, you can actually act on that instead of paying frontier prices for everything because migrating is a sprint.

Concretely: a 2,000-token prompt with a 500-token answer on gpt-6-sol costs about 7.6 credits — under one US cent.

Billing is prepaid and settles on real usage. A text request reserves its worst-case cost up front, then charges what the response actually consumed and returns the difference. There is no monthly commitment and no per-day request quota; the only gate is whether your balance covers the call.

The same key does images, speech and video

# Image
curl -X POST https://router.qtum.ai/v1/images/generations \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-image-2","prompt":"a clean product render of a red apple on a white table","size":"1024x1024","n":1}'

# Text to speech
curl -X POST https://router.qtum.ai/v1/audio/speech \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-tts-flash","input":"Hello from the Qtum router.","voice":"Cherry","response_format":"mp3"}' \
  --output speech.mp3

# Video (async: create a task, then poll it)
curl -X POST https://router.qtum.ai/v1/videos \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: video-job-001" \
  -d '{"model":"doubao-seedance-2-0-260128","prompt":"a calm ocean at sunset, cinematic","resolution":"720p","seconds":5,"mode":"text-to-video"}'

Same host, same key, same balance. The integration guide has the Python and TypeScript versions of each of these, plus the polling loop for video and the reference-file upload for image-to-video.

If you would rather not write any code

Qtum Video is the same video models behind a prompt box in the browser — no key, no SDK, no polling loop. Ten models, text-to-video and image-to-video, up to 30 seconds depending on the model.

Video is priced per five seconds of output, and the arithmetic is public: ceil(credits_per_5s × seconds ÷ 5). So at 720p:

ModelPer 5sA 10-second clip
doubao-seedance-1-5-pro-251215 (480p)80 credits160 credits ≈ $0.16
wan3.0-video347 credits694 credits ≈ $0.69
doubao-seedance-2-0-260128994 credits1,988 credits ≈ $1.99

Two things worth knowing before you spend anything. A failed generation is refunded in full — credits are reserved when the task is created and returned if it does not produce a video, so a provider error costs you time and not money. And there is no daily or monthly generation limit; the only check is your balance. The full price table lists every model, every resolution and every clip length.

What it costs to find out

Nothing, to look: the model catalogue and the video prices are public pages, no account needed. When you want to make a call, the Router page is where you get a key — then top up and point your existing code at the base URL above. The first request should take about as long as reading this paragraph.

And if you only remember one thing: your model choice should be a config value. Everything above is in service of that.

One API key for every model

The Qtum AI Router puts text, image, audio and video models behind a single OpenAI- and Anthropic-compatible endpoint, billed in one place.