Integration guide

The Qtum AI Router speaks the OpenAI Chat Completions API and the Anthropic Messages API. If your code already calls either one, two values change: the base URL and the key.

Quick start

Base URL: https://router.qtum.ai/v1. Authenticate with a Qtum API key created in the console.

cURL
curl https://router.qtum.ai/v1/chat/completions \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.5",
    "messages": [
      { "role": "user", "content": "Explain Qtum AAL in one paragraph." }
    ]
  }'
Python · openai SDK
from openai import OpenAI

client = OpenAI(
    base_url="https://router.qtum.ai/v1",
    api_key="$QTUM_API_KEY",
)

resp = client.chat.completions.create(
    model="gpt-5.5",
    messages=[{"role": "user", "content": "Explain Qtum AAL in one paragraph."}],
)
print(resp.choices[0].message.content)
Node · openai SDK
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://router.qtum.ai/v1",
  apiKey: process.env.QTUM_API_KEY,
});

const resp = await client.chat.completions.create({
  model: "gpt-5.5",
  messages: [{ role: "user", content: "Explain Qtum AAL in one paragraph." }],
});
console.log(resp.choices[0].message.content);
Environment variables
# Drop-in replacement: change two env vars, keep your existing code.
export OPENAI_BASE_URL=https://router.qtum.ai/v1
export OPENAI_API_KEY=$QTUM_API_KEY

# Then any OpenAI-compatible SDK or tool just works.

Model ids come from the Models page.

Authentication

Send your key as a bearer token. Where a client only supports the Anthropic style, x-api-key works too.

Headers
Authorization: Bearer $QTUM_API_KEY
Content-Type: application/json

Text & vision

Two request shapes reach the same models. Use whichever your client already speaks.

OpenAI — POST /v1/chat/completions

cURL
curl https://router.qtum.ai/v1/chat/completions \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.5",
    "messages": [
      { "role": "user", "content": "Explain Qtum AAL in one paragraph." }
    ]
  }'
Python
from openai import OpenAI

client = OpenAI(
    base_url="https://router.qtum.ai/v1",
    api_key="$QTUM_API_KEY",
)

resp = client.chat.completions.create(
    model="gpt-5.5",
    messages=[{"role": "user", "content": "Explain Qtum AAL in one paragraph."}],
)
print(resp.choices[0].message.content)

Anthropic — POST https://router.qtum.ai/v1/messages

cURL
curl https://router.qtum.ai/v1/messages \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4-7",
    "max_tokens": 1024,
    "messages": [
      { "role": "user", "content": "Explain Qtum AAL in one paragraph." }
    ]
  }'
Python · anthropic SDK
import anthropic

# SDK appends /v1/messages itself — base_url stops at the host.
client = anthropic.Anthropic(
    base_url="https://router.qtum.ai",
    api_key="$QTUM_API_KEY",
)

msg = client.messages.create(
    model="claude-opus-4-7",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Explain Qtum AAL in one paragraph."}],
)
print(msg.content[0].text)
Environment variables
# Drop-in for Claude Code / claude-code-router (Bearer auth):
export ANTHROPIC_BASE_URL=https://router.qtum.ai
export ANTHROPIC_AUTH_TOKEN=$QTUM_API_KEY

# For the anthropic Python / JS SDK, use ANTHROPIC_API_KEY instead (x-api-key auth).

Streaming

Streaming is standard server-sent events, so existing stream handling works unchanged. The usage reported in the stream is what the request settles against.

Python
from openai import OpenAI

client = OpenAI(base_url="https://router.qtum.ai/v1", api_key="$QTUM_API_KEY")

stream = client.chat.completions.create(
    model="gpt-5.5",
    messages=[{"role": "user", "content": "Tell me a story."}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content or ""
    print(delta, end="", flush=True)

Images

Generation and editing, in the OpenAI shapes.

Generate · cURL
curl -X POST https://router.qtum.ai/v1/images/generations \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-image-2-2026-04-21",
    "prompt": "a clean product render of a red apple on a white table",
    "size": "1024x1024",
    "n": 1
  }'
Generate · Python
import base64
import requests
from openai import OpenAI

client = OpenAI(base_url="https://router.qtum.ai/v1", api_key="$QTUM_API_KEY")

img = client.images.generate(
    model="gpt-image-2-2026-04-21",
    prompt="a clean product render of a red apple on a white table",
    size="1024x1024",
    n=1,
)

# OpenAI image models commonly return b64_json; Qwen/Wan commonly return a URL.
item = img.data[0]
if item.url:
    image_bytes = requests.get(item.url, timeout=60).content
elif item.b64_json:
    image_bytes = base64.b64decode(item.b64_json)
else:
    raise ValueError("image response missing url and b64_json")

with open("output.png", "wb") as f:
    f.write(image_bytes)
Edit · cURL
curl -X POST https://router.qtum.ai/v1/images/edits \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -F "model=gpt-image-2-2026-04-21" \
  -F "prompt=replace the background with a clean white studio" \
  -F "image=@./input.png" \
  -F "size=1024x1024" \
  -F "n=1"

Reading the response

Model familyFieldHow to handle it
OpenAI image modelsdata[].b64_jsonBase64-decode the value and save the resulting image bytes.
Qwen / Wan image modelsdata[].urlDownload the temporary URL promptly; it is commonly valid for about 24 hours.

Audio

Text-to-speech on /v1/audio/speech, transcription on /v1/audio/transcriptions.

Text to speech · cURL
curl -X POST https://router.qtum.ai/v1/audio/speech \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-tts-flash",
    "input": "Hello from the Qtum router.",
    "voice": "Cherry",
    "response_format": "mp3"
  }' \
  --output speech.mp3
Text to speech · Python
from openai import OpenAI

client = OpenAI(base_url="https://router.qtum.ai/v1", api_key="$QTUM_API_KEY")

with client.audio.speech.with_streaming_response.create(
    model="qwen3-tts-flash",
    voice="Cherry",
    input="Hello from the Qtum router.",
    response_format="mp3",
) as resp:
    resp.stream_to_file("speech.mp3")
Transcription · cURL
curl -X POST https://router.qtum.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -F "model=qwen3-asr-flash" \
  -F "language=en" \
  -F "response_format=json" \
  -F "file=@./speech.mp3"
Transcription · Python
from openai import OpenAI

client = OpenAI(base_url="https://router.qtum.ai/v1", api_key="$QTUM_API_KEY")

with open("speech.mp3", "rb") as f:
    tr = client.audio.transcriptions.create(
        model="qwen3-asr-flash",
        file=f,
        language="en",
        response_format="json",
    )
print(tr.text)

Text-to-speech models

FamilyModelFormats & voicesBilled on
Qwen text-to-speechqwen3-tts-flashQwen voice names; output formats supported by the modelCharacters
OpenAI text-to-speechgpt-4o-mini-tts-2025-12-15OpenAI-compatible voices and response formatsTokens

Transcription models

FamilyModelFormats & voicesBilled on
Qwen speech-to-textqwen3-asr-flashjson, textAudio seconds
Whisperwhisper-1json, text, srt, vtt, verbose_jsonAudio seconds
GPT speech-to-textgpt-4o-transcribejson, textTokens

Video

Video is asynchronous: create a task, poll its status, then fetch the content.

cURL
# 1) Text-to-video
curl -X POST https://router.qtum.ai/v1/videos \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: video-job-001" \
  -d '{
    "model": "doubao-seedance-2-0-260128",
    "prompt": "a calm ocean at sunset, cinematic",
    "resolution": "720p",
    "seconds": 5,
    "mode": "text-to-video"
  }'
# → {"id":"<id>","object":"video","status":"queued"}

# 2) Poll every 2–5 seconds until completed, failed, or cancelled
curl https://router.qtum.ai/v1/videos/<id> \
  -H "Authorization: Bearer $QTUM_API_KEY"

# 3) Download the completed result; the gateway streams the video bytes
curl https://router.qtum.ai/v1/videos/<id>/content \
  -H "Authorization: Bearer $QTUM_API_KEY" -o out.mp4

# --- Image-to-video ---
# Upload the reference once. Images, videos, and MP3/WAV reference audio are accepted;
# purpose is forced to video.
curl https://router.qtum.ai/v1/files \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -F "[email protected]"
# → {"id":"file-aaaa","object":"file","purpose":"video",...}

curl -X POST https://router.qtum.ai/v1/videos \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: image-video-job-001" \
  -d '{
    "model": "doubao-seedance-2-0-260128",
    "prompt": "the reference photo comes alive, gentle push-in",
    "resolution": "720p",
    "seconds": 5,
    "mode": "image-to-video",
    "input_reference": { "type": "image", "url": "file-aaaa" }
  }'

# Multiple references use content[] with one uploaded file id per image_url.url.
# Give each image role=reference_image.

# --- Video-to-video ---
curl https://router.qtum.ai/v1/files \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -F "[email protected]"
# → {"id":"file-bbbb","object":"file","purpose":"video",...}

curl -X POST https://router.qtum.ai/v1/videos \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: video-video-job-001" \
  -d '{
    "model": "doubao-seedance-2-0-260128",
    "prompt": "extend the reference video with a gentle rain effect",
    "resolution": "720p",
    "seconds": 5,
    "mode": "video-to-video",
    "input_reference": { "type": "video", "url": "file-bbbb", "duration": 7.4 }
  }'

# --- Seedance 2.5 reference audio ---
# The upload endpoint accepts MP3/WAV reference audio and returns a reusable file id.
curl https://router.qtum.ai/v1/files \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -F "[email protected]"
# → {"id":"file-cccc","object":"file","purpose":"video",...}

# Reference audio belongs in content[].audio_url, not input_reference.
curl -X POST https://router.qtum.ai/v1/videos \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: seedance-audio-job-001" \
  -d '{
    "model": "doubao-seedance-2-5-260628",
    "content": [
      { "type": "text", "text": "cut the scene to the rhythm of audio 1" },
      {
        "type": "audio_url",
        "audio_url": { "url": "file-cccc" },
        "role": "reference_audio"
      }
    ],
    "resolution": "720p",
    "seconds": 5
  }'

# --- HappyHorse Video Edit (special case) ---
# The model derives output duration from the input video. Do not send a top-level
# seconds or duration field. input_reference.duration is required (whole seconds, 3–60).
curl -X POST https://router.qtum.ai/v1/videos \
  -H "Authorization: Bearer $QTUM_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: happyhorse-video-edit-001" \
  -d '{
    "model": "happyhorse-1.0-video-edit",
    "prompt": "increase the warm lighting while preserving the scene",
    "resolution": "720p",
    "mode": "video-edit",
    "input_reference": { "type": "video", "url": "file-bbbb", "duration": 7 }
  }'
Python
import time
import requests

BASE = "https://router.qtum.ai/v1"
AUTH = {"Authorization": "Bearer $QTUM_API_KEY"}

# Create a text-to-video task. Use a new idempotency key for each distinct job.
create_headers = {**AUTH, "Idempotency-Key": "video-job-001"}
vid = requests.post(f"{BASE}/videos", headers=create_headers, json={
    "model": "doubao-seedance-2-0-260128",
    "prompt": "a calm ocean at sunset, cinematic",
    "resolution": "720p",
    "seconds": 5,
    "mode": "text-to-video",
}).json()["id"]

# Poll with gradual backoff until a terminal state.
delay = 3
while True:
    task = requests.get(f"{BASE}/videos/{vid}", headers=AUTH).json()
    if task["status"] in ("completed", "failed", "cancelled"):
        break
    time.sleep(delay)
    delay = min(delay * 1.5, 30)

# The content endpoint returns the video body directly after completion.
if task["status"] == "completed":
    with requests.get(f"{BASE}/videos/{vid}/content", headers=AUTH, stream=True) as resp:
        resp.raise_for_status()
        with open("out.mp4", "wb") as f:
            for chunk in resp.iter_content(1024 * 1024):
                f.write(chunk)

# Upload reference media once, then reuse the returned file id in the video request.
with open("ref1.png", "rb") as f:
    image_fid = requests.post(f"{BASE}/files", headers=AUTH,
                              files={"file": f}).json()["id"]

image_job = requests.post(f"{BASE}/videos", headers={
    **AUTH, "Idempotency-Key": "image-video-job-001"
}, json={
    "model": "doubao-seedance-2-0-260128",
    "prompt": "the reference photo comes alive, gentle push-in",
    "resolution": "720p",
    "seconds": 5,
    "mode": "image-to-video",
    "input_reference": {"type": "image", "url": image_fid},
}).json()["id"]

with open("source.mp4", "rb") as f:
    video_fid = requests.post(f"{BASE}/files", headers=AUTH,
                              files={"file": f}).json()["id"]

video_job = requests.post(f"{BASE}/videos", headers={
    **AUTH, "Idempotency-Key": "video-video-job-001"
}, json={
    "model": "doubao-seedance-2-0-260128",
    "prompt": "extend the reference video with a gentle rain effect",
    "resolution": "720p",
    "seconds": 5,
    "mode": "video-to-video",
    "input_reference": {"type": "video", "url": video_fid, "duration": 7.4},
}).json()["id"]

with open("soundtrack.mp3", "rb") as f:
    audio_fid = requests.post(f"{BASE}/files", headers=AUTH,
                              files={"file": f}).json()["id"]

audio_reference_job = requests.post(f"{BASE}/videos", headers={
    **AUTH, "Idempotency-Key": "seedance-audio-job-001"
}, json={
    "model": "doubao-seedance-2-5-260628",
    "content": [
        {"type": "text", "text": "cut the scene to the rhythm of audio 1"},
        {
            "type": "audio_url",
            "audio_url": {"url": audio_fid},
            "role": "reference_audio",
        },
    ],
    "resolution": "720p",
    "seconds": 5,
}).json()["id"]

# HappyHorse Video Edit derives output duration from the input video. Do not
# include top-level seconds or duration. The reference duration is required.
edit_job = requests.post(f"{BASE}/videos", headers={
    **AUTH, "Idempotency-Key": "happyhorse-video-edit-001"
}, json={
    "model": "happyhorse-1.0-video-edit",
    "prompt": "increase the warm lighting while preserving the scene",
    "resolution": "720p",
    "mode": "video-edit",
    "input_reference": {"type": "video", "url": video_fid, "duration": 7},
}).json()["id"]
Task response
{
  "id": "b36ad028-1414-4bf5-8975-c7a66bd330a9",
  "object": "video",
  "model": "doubao-seedance-2-0-260128",
  "status": "queued",
  "created_at": 1718500000,
  "completed_at": null,
  "result_file_id": null,
  "seconds": null,
  "usage": null,
  "error": null
}

Request fields

FieldTypeRequiredDescription
modelstringRequiredAn active video model ID from the Models page.
promptstringConditionalText instruction. Required unless the model receives equivalent text in content[].
contentarrayConditionalMultimodal input used by Seedance and native MiniMax-H3 requests. Seedance reference audio uses an audio_url item with role=reference_audio.
resolutionstringModel-specificSeedance, Wan, and HappyHorse use their supported 480p/720p/1080p tiers. HappyHorse Video Edit supports only 720p or 1080p. MiniMax-H3 uses 768P or 2K.
sizestringSoraSora frame size, for example 1280x720. Use this instead of resolution.
secondsintegerModel-specificRequested output duration. Accepted values are determined by the selected model. Do not send this field for happyhorse-1.0-video-edit.
durationintegerMiniMax-H3MiniMax-H3 output duration from 4 to 15 seconds; seconds is also accepted as an alias.
modestringOptionalGeneration mode such as text-to-video, image-to-video, video-to-video, multimodal-reference, or video-edit. Supported values are model-specific.
input_referenceobjectConditionalA single reference with type and url. A video reference must also include duration; HappyHorse Video Edit requires a video reference with a whole duration from 3 to 60 seconds.
request_idstringOptionalForwarded provider request identifier. Prefer the standard Idempotency-Key header for safe retries.

Response fields

FieldTypeDescription
idstringVideo task ID used for status polling and content download.
objectstringThe normalized object type: video.
modelstringModel ID used by this task.
statusstringqueued, in_progress, completed, failed, or cancelled.
created_atintegerCreation time as a Unix timestamp in seconds.
completed_atinteger | nullCompletion time, or null before completion.
result_file_idstring | nullProvider result file ID when supplied; clients can always use the content endpoint.
secondsnumber | nullActual output seconds for second-billed models; may be null while pending or for token-billed models.
usageobject | nullToken usage for token-billed video models when supplied by the provider.
errorobject | nullFailure details containing code and message.

Task statuses

StatusMeaningWhat to do
queuedTask accepted and waiting for provider processing.Show queued and continue polling.
in_progressThe provider is generating the video.Show generating and continue polling.
completedGeneration finished successfully.Fetch /v1/videos/{id}/content.
failedGeneration failed.Read error.code and error.message.
cancelledThe task was cancelled.Stop polling and show the cancelled state.

Per-model behaviour

FamilyFields it usesDurationBilled on
Doubao Seedanceresolution, seconds, mode, input_reference / contentModel-specific; upstream validates the valueCompletion tokens by resolution tier
Sorasize, prompt, optional input_referencesora-2 uses its 4s default; Pro/Azure accept advertised durationsOutput seconds by resolution tier
MiniMax-H3prompt or content, resolution, duration, optional ratio/mediaWhole seconds from 4 to 15; default 4Output seconds at 768p or 2K rate
HappyHorse Video Editresolution, prompt, mode=video-edit, input_reference videoDerived from input video; require input_reference.duration (3–60 whole seconds); omit top-level seconds/durationActual output seconds at 720p or 1080p rate; completed seconds may be fractional

Video errors

HTTPCodeMeaning
400invalid_paramsMalformed input, unsupported duration, or unsupported resolution.
400invalid_modelThe model is not active or is not available for video generation.
400upstream video errorsProvider validation such as input_video_too_long or video_resolution_not_supported is forwarded unchanged.
402insufficient_balanceThe Qtum account does not meet the minimum credit balance.
404not_foundThe task does not exist or does not belong to the current user.
409idempotency_conflictThe request could not be safely deduplicated; retry with a new Idempotency-Key.
502/504upstream_error / upstream_timeoutThe video provider is unavailable or did not accept the request before timeout.
503video_unavailableRouter video generation is not configured or is temporarily unavailable.

Pricing & credits

Usage is billed in credits, where 1000 credits equal 1 USD of deposit.

A request reserves worst-case credits before it runs, settles against the usage the model actually reported, and refunds the difference. Failed requests are not billed — credits settle only on a successful response. Current per-model rates are on the Models page.

Listing models

The same catalog the Models page shows, as JSON.

cURL
curl https://router.qtum.ai/v1/models -H "Authorization: Bearer $QTUM_API_KEY"

Error codes

Errors return { "error": { "code", "message" } }. The HTTP status and the stable code string are below. The Anthropic endpoint returns Anthropic's native { "type": "error", "error": { "type", "message" } } envelope instead.

HTTPCodeMeaning
400invalid_paramsMalformed body, or a missing / invalid model field
400invalid_modelModel id is not available for that endpoint
401unauthorizedMissing, invalid, or revoked API key
402insufficient_balanceInsufficient credits — top up to continue
404not_foundVideo task id not found (video status / content polling)
500internal_errorInternal error (e.g. pricing misconfigured) — not billed
502upstream_errorThe gateway could not reach the model provider — not billed
503video_unavailableVideo generation is temporarily unavailable
504upstream_timeoutThe gateway timed out reaching the provider (e.g. an input image too large to upload in time) — not billed

Ready to call it

Create a key in the console, then point your client at https://router.qtum.ai/v1.