Integration guide
The Qtum AI Router speaks the OpenAI Chat Completions API and the Anthropic Messages API. If your code already calls either one, two values change: the base URL and the key.
Quick start
Base URL: https://router.qtum.ai/v1. Authenticate with a Qtum API key created in the console.
curl https://router.qtum.ai/v1/chat/completions \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.5",
"messages": [
{ "role": "user", "content": "Explain Qtum AAL in one paragraph." }
]
}'from openai import OpenAI
client = OpenAI(
base_url="https://router.qtum.ai/v1",
api_key="$QTUM_API_KEY",
)
resp = client.chat.completions.create(
model="gpt-5.5",
messages=[{"role": "user", "content": "Explain Qtum AAL in one paragraph."}],
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://router.qtum.ai/v1",
apiKey: process.env.QTUM_API_KEY,
});
const resp = await client.chat.completions.create({
model: "gpt-5.5",
messages: [{ role: "user", content: "Explain Qtum AAL in one paragraph." }],
});
console.log(resp.choices[0].message.content);# Drop-in replacement: change two env vars, keep your existing code.
export OPENAI_BASE_URL=https://router.qtum.ai/v1
export OPENAI_API_KEY=$QTUM_API_KEY
# Then any OpenAI-compatible SDK or tool just works.Model ids come from the Models page.
Authentication
Send your key as a bearer token. Where a client only supports the Anthropic style, x-api-key works too.
Authorization: Bearer $QTUM_API_KEY
Content-Type: application/jsonText & vision
Two request shapes reach the same models. Use whichever your client already speaks.
OpenAI — POST /v1/chat/completions
curl https://router.qtum.ai/v1/chat/completions \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.5",
"messages": [
{ "role": "user", "content": "Explain Qtum AAL in one paragraph." }
]
}'from openai import OpenAI
client = OpenAI(
base_url="https://router.qtum.ai/v1",
api_key="$QTUM_API_KEY",
)
resp = client.chat.completions.create(
model="gpt-5.5",
messages=[{"role": "user", "content": "Explain Qtum AAL in one paragraph."}],
)
print(resp.choices[0].message.content)Anthropic — POST https://router.qtum.ai/v1/messages
curl https://router.qtum.ai/v1/messages \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4-7",
"max_tokens": 1024,
"messages": [
{ "role": "user", "content": "Explain Qtum AAL in one paragraph." }
]
}'import anthropic
# SDK appends /v1/messages itself — base_url stops at the host.
client = anthropic.Anthropic(
base_url="https://router.qtum.ai",
api_key="$QTUM_API_KEY",
)
msg = client.messages.create(
model="claude-opus-4-7",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain Qtum AAL in one paragraph."}],
)
print(msg.content[0].text)# Drop-in for Claude Code / claude-code-router (Bearer auth):
export ANTHROPIC_BASE_URL=https://router.qtum.ai
export ANTHROPIC_AUTH_TOKEN=$QTUM_API_KEY
# For the anthropic Python / JS SDK, use ANTHROPIC_API_KEY instead (x-api-key auth).Streaming
Streaming is standard server-sent events, so existing stream handling works unchanged. The usage reported in the stream is what the request settles against.
from openai import OpenAI
client = OpenAI(base_url="https://router.qtum.ai/v1", api_key="$QTUM_API_KEY")
stream = client.chat.completions.create(
model="gpt-5.5",
messages=[{"role": "user", "content": "Tell me a story."}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content or ""
print(delta, end="", flush=True)Images
Generation and editing, in the OpenAI shapes.
curl -X POST https://router.qtum.ai/v1/images/generations \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-2-2026-04-21",
"prompt": "a clean product render of a red apple on a white table",
"size": "1024x1024",
"n": 1
}'import base64
import requests
from openai import OpenAI
client = OpenAI(base_url="https://router.qtum.ai/v1", api_key="$QTUM_API_KEY")
img = client.images.generate(
model="gpt-image-2-2026-04-21",
prompt="a clean product render of a red apple on a white table",
size="1024x1024",
n=1,
)
# OpenAI image models commonly return b64_json; Qwen/Wan commonly return a URL.
item = img.data[0]
if item.url:
image_bytes = requests.get(item.url, timeout=60).content
elif item.b64_json:
image_bytes = base64.b64decode(item.b64_json)
else:
raise ValueError("image response missing url and b64_json")
with open("output.png", "wb") as f:
f.write(image_bytes)curl -X POST https://router.qtum.ai/v1/images/edits \
-H "Authorization: Bearer $QTUM_API_KEY" \
-F "model=gpt-image-2-2026-04-21" \
-F "prompt=replace the background with a clean white studio" \
-F "image=@./input.png" \
-F "size=1024x1024" \
-F "n=1"Reading the response
| Model family | Field | How to handle it |
|---|---|---|
| OpenAI image models | data[].b64_json | Base64-decode the value and save the resulting image bytes. |
| Qwen / Wan image models | data[].url | Download the temporary URL promptly; it is commonly valid for about 24 hours. |
Audio
Text-to-speech on /v1/audio/speech, transcription on /v1/audio/transcriptions.
curl -X POST https://router.qtum.ai/v1/audio/speech \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-tts-flash",
"input": "Hello from the Qtum router.",
"voice": "Cherry",
"response_format": "mp3"
}' \
--output speech.mp3from openai import OpenAI
client = OpenAI(base_url="https://router.qtum.ai/v1", api_key="$QTUM_API_KEY")
with client.audio.speech.with_streaming_response.create(
model="qwen3-tts-flash",
voice="Cherry",
input="Hello from the Qtum router.",
response_format="mp3",
) as resp:
resp.stream_to_file("speech.mp3")curl -X POST https://router.qtum.ai/v1/audio/transcriptions \
-H "Authorization: Bearer $QTUM_API_KEY" \
-F "model=qwen3-asr-flash" \
-F "language=en" \
-F "response_format=json" \
-F "file=@./speech.mp3"from openai import OpenAI
client = OpenAI(base_url="https://router.qtum.ai/v1", api_key="$QTUM_API_KEY")
with open("speech.mp3", "rb") as f:
tr = client.audio.transcriptions.create(
model="qwen3-asr-flash",
file=f,
language="en",
response_format="json",
)
print(tr.text)Text-to-speech models
| Family | Model | Formats & voices | Billed on |
|---|---|---|---|
| Qwen text-to-speech | qwen3-tts-flash | Qwen voice names; output formats supported by the model | Characters |
| OpenAI text-to-speech | gpt-4o-mini-tts-2025-12-15 | OpenAI-compatible voices and response formats | Tokens |
Transcription models
| Family | Model | Formats & voices | Billed on |
|---|---|---|---|
| Qwen speech-to-text | qwen3-asr-flash | json, text | Audio seconds |
| Whisper | whisper-1 | json, text, srt, vtt, verbose_json | Audio seconds |
| GPT speech-to-text | gpt-4o-transcribe | json, text | Tokens |
Video
Video is asynchronous: create a task, poll its status, then fetch the content.
# 1) Text-to-video
curl -X POST https://router.qtum.ai/v1/videos \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-job-001" \
-d '{
"model": "doubao-seedance-2-0-260128",
"prompt": "a calm ocean at sunset, cinematic",
"resolution": "720p",
"seconds": 5,
"mode": "text-to-video"
}'
# → {"id":"<id>","object":"video","status":"queued"}
# 2) Poll every 2–5 seconds until completed, failed, or cancelled
curl https://router.qtum.ai/v1/videos/<id> \
-H "Authorization: Bearer $QTUM_API_KEY"
# 3) Download the completed result; the gateway streams the video bytes
curl https://router.qtum.ai/v1/videos/<id>/content \
-H "Authorization: Bearer $QTUM_API_KEY" -o out.mp4
# --- Image-to-video ---
# Upload the reference once. Images, videos, and MP3/WAV reference audio are accepted;
# purpose is forced to video.
curl https://router.qtum.ai/v1/files \
-H "Authorization: Bearer $QTUM_API_KEY" \
-F "[email protected]"
# → {"id":"file-aaaa","object":"file","purpose":"video",...}
curl -X POST https://router.qtum.ai/v1/videos \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: image-video-job-001" \
-d '{
"model": "doubao-seedance-2-0-260128",
"prompt": "the reference photo comes alive, gentle push-in",
"resolution": "720p",
"seconds": 5,
"mode": "image-to-video",
"input_reference": { "type": "image", "url": "file-aaaa" }
}'
# Multiple references use content[] with one uploaded file id per image_url.url.
# Give each image role=reference_image.
# --- Video-to-video ---
curl https://router.qtum.ai/v1/files \
-H "Authorization: Bearer $QTUM_API_KEY" \
-F "[email protected]"
# → {"id":"file-bbbb","object":"file","purpose":"video",...}
curl -X POST https://router.qtum.ai/v1/videos \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-video-job-001" \
-d '{
"model": "doubao-seedance-2-0-260128",
"prompt": "extend the reference video with a gentle rain effect",
"resolution": "720p",
"seconds": 5,
"mode": "video-to-video",
"input_reference": { "type": "video", "url": "file-bbbb", "duration": 7.4 }
}'
# --- Seedance 2.5 reference audio ---
# The upload endpoint accepts MP3/WAV reference audio and returns a reusable file id.
curl https://router.qtum.ai/v1/files \
-H "Authorization: Bearer $QTUM_API_KEY" \
-F "[email protected]"
# → {"id":"file-cccc","object":"file","purpose":"video",...}
# Reference audio belongs in content[].audio_url, not input_reference.
curl -X POST https://router.qtum.ai/v1/videos \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: seedance-audio-job-001" \
-d '{
"model": "doubao-seedance-2-5-260628",
"content": [
{ "type": "text", "text": "cut the scene to the rhythm of audio 1" },
{
"type": "audio_url",
"audio_url": { "url": "file-cccc" },
"role": "reference_audio"
}
],
"resolution": "720p",
"seconds": 5
}'
# --- HappyHorse Video Edit (special case) ---
# The model derives output duration from the input video. Do not send a top-level
# seconds or duration field. input_reference.duration is required (whole seconds, 3–60).
curl -X POST https://router.qtum.ai/v1/videos \
-H "Authorization: Bearer $QTUM_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: happyhorse-video-edit-001" \
-d '{
"model": "happyhorse-1.0-video-edit",
"prompt": "increase the warm lighting while preserving the scene",
"resolution": "720p",
"mode": "video-edit",
"input_reference": { "type": "video", "url": "file-bbbb", "duration": 7 }
}'import time
import requests
BASE = "https://router.qtum.ai/v1"
AUTH = {"Authorization": "Bearer $QTUM_API_KEY"}
# Create a text-to-video task. Use a new idempotency key for each distinct job.
create_headers = {**AUTH, "Idempotency-Key": "video-job-001"}
vid = requests.post(f"{BASE}/videos", headers=create_headers, json={
"model": "doubao-seedance-2-0-260128",
"prompt": "a calm ocean at sunset, cinematic",
"resolution": "720p",
"seconds": 5,
"mode": "text-to-video",
}).json()["id"]
# Poll with gradual backoff until a terminal state.
delay = 3
while True:
task = requests.get(f"{BASE}/videos/{vid}", headers=AUTH).json()
if task["status"] in ("completed", "failed", "cancelled"):
break
time.sleep(delay)
delay = min(delay * 1.5, 30)
# The content endpoint returns the video body directly after completion.
if task["status"] == "completed":
with requests.get(f"{BASE}/videos/{vid}/content", headers=AUTH, stream=True) as resp:
resp.raise_for_status()
with open("out.mp4", "wb") as f:
for chunk in resp.iter_content(1024 * 1024):
f.write(chunk)
# Upload reference media once, then reuse the returned file id in the video request.
with open("ref1.png", "rb") as f:
image_fid = requests.post(f"{BASE}/files", headers=AUTH,
files={"file": f}).json()["id"]
image_job = requests.post(f"{BASE}/videos", headers={
**AUTH, "Idempotency-Key": "image-video-job-001"
}, json={
"model": "doubao-seedance-2-0-260128",
"prompt": "the reference photo comes alive, gentle push-in",
"resolution": "720p",
"seconds": 5,
"mode": "image-to-video",
"input_reference": {"type": "image", "url": image_fid},
}).json()["id"]
with open("source.mp4", "rb") as f:
video_fid = requests.post(f"{BASE}/files", headers=AUTH,
files={"file": f}).json()["id"]
video_job = requests.post(f"{BASE}/videos", headers={
**AUTH, "Idempotency-Key": "video-video-job-001"
}, json={
"model": "doubao-seedance-2-0-260128",
"prompt": "extend the reference video with a gentle rain effect",
"resolution": "720p",
"seconds": 5,
"mode": "video-to-video",
"input_reference": {"type": "video", "url": video_fid, "duration": 7.4},
}).json()["id"]
with open("soundtrack.mp3", "rb") as f:
audio_fid = requests.post(f"{BASE}/files", headers=AUTH,
files={"file": f}).json()["id"]
audio_reference_job = requests.post(f"{BASE}/videos", headers={
**AUTH, "Idempotency-Key": "seedance-audio-job-001"
}, json={
"model": "doubao-seedance-2-5-260628",
"content": [
{"type": "text", "text": "cut the scene to the rhythm of audio 1"},
{
"type": "audio_url",
"audio_url": {"url": audio_fid},
"role": "reference_audio",
},
],
"resolution": "720p",
"seconds": 5,
}).json()["id"]
# HappyHorse Video Edit derives output duration from the input video. Do not
# include top-level seconds or duration. The reference duration is required.
edit_job = requests.post(f"{BASE}/videos", headers={
**AUTH, "Idempotency-Key": "happyhorse-video-edit-001"
}, json={
"model": "happyhorse-1.0-video-edit",
"prompt": "increase the warm lighting while preserving the scene",
"resolution": "720p",
"mode": "video-edit",
"input_reference": {"type": "video", "url": video_fid, "duration": 7},
}).json()["id"]{
"id": "b36ad028-1414-4bf5-8975-c7a66bd330a9",
"object": "video",
"model": "doubao-seedance-2-0-260128",
"status": "queued",
"created_at": 1718500000,
"completed_at": null,
"result_file_id": null,
"seconds": null,
"usage": null,
"error": null
}Request fields
| Field | Type | Required | Description |
|---|---|---|---|
| model | string | Required | An active video model ID from the Models page. |
| prompt | string | Conditional | Text instruction. Required unless the model receives equivalent text in content[]. |
| content | array | Conditional | Multimodal input used by Seedance and native MiniMax-H3 requests. Seedance reference audio uses an audio_url item with role=reference_audio. |
| resolution | string | Model-specific | Seedance, Wan, and HappyHorse use their supported 480p/720p/1080p tiers. HappyHorse Video Edit supports only 720p or 1080p. MiniMax-H3 uses 768P or 2K. |
| size | string | Sora | Sora frame size, for example 1280x720. Use this instead of resolution. |
| seconds | integer | Model-specific | Requested output duration. Accepted values are determined by the selected model. Do not send this field for happyhorse-1.0-video-edit. |
| duration | integer | MiniMax-H3 | MiniMax-H3 output duration from 4 to 15 seconds; seconds is also accepted as an alias. |
| mode | string | Optional | Generation mode such as text-to-video, image-to-video, video-to-video, multimodal-reference, or video-edit. Supported values are model-specific. |
| input_reference | object | Conditional | A single reference with type and url. A video reference must also include duration; HappyHorse Video Edit requires a video reference with a whole duration from 3 to 60 seconds. |
| request_id | string | Optional | Forwarded provider request identifier. Prefer the standard Idempotency-Key header for safe retries. |
Response fields
| Field | Type | Description |
|---|---|---|
| id | string | Video task ID used for status polling and content download. |
| object | string | The normalized object type: video. |
| model | string | Model ID used by this task. |
| status | string | queued, in_progress, completed, failed, or cancelled. |
| created_at | integer | Creation time as a Unix timestamp in seconds. |
| completed_at | integer | null | Completion time, or null before completion. |
| result_file_id | string | null | Provider result file ID when supplied; clients can always use the content endpoint. |
| seconds | number | null | Actual output seconds for second-billed models; may be null while pending or for token-billed models. |
| usage | object | null | Token usage for token-billed video models when supplied by the provider. |
| error | object | null | Failure details containing code and message. |
Task statuses
| Status | Meaning | What to do |
|---|---|---|
| queued | Task accepted and waiting for provider processing. | Show queued and continue polling. |
| in_progress | The provider is generating the video. | Show generating and continue polling. |
| completed | Generation finished successfully. | Fetch /v1/videos/{id}/content. |
| failed | Generation failed. | Read error.code and error.message. |
| cancelled | The task was cancelled. | Stop polling and show the cancelled state. |
Per-model behaviour
| Family | Fields it uses | Duration | Billed on |
|---|---|---|---|
| Doubao Seedance | resolution, seconds, mode, input_reference / content | Model-specific; upstream validates the value | Completion tokens by resolution tier |
| Sora | size, prompt, optional input_reference | sora-2 uses its 4s default; Pro/Azure accept advertised durations | Output seconds by resolution tier |
| MiniMax-H3 | prompt or content, resolution, duration, optional ratio/media | Whole seconds from 4 to 15; default 4 | Output seconds at 768p or 2K rate |
| HappyHorse Video Edit | resolution, prompt, mode=video-edit, input_reference video | Derived from input video; require input_reference.duration (3–60 whole seconds); omit top-level seconds/duration | Actual output seconds at 720p or 1080p rate; completed seconds may be fractional |
Video errors
| HTTP | Code | Meaning |
|---|---|---|
| 400 | invalid_params | Malformed input, unsupported duration, or unsupported resolution. |
| 400 | invalid_model | The model is not active or is not available for video generation. |
| 400 | upstream video errors | Provider validation such as input_video_too_long or video_resolution_not_supported is forwarded unchanged. |
| 402 | insufficient_balance | The Qtum account does not meet the minimum credit balance. |
| 404 | not_found | The task does not exist or does not belong to the current user. |
| 409 | idempotency_conflict | The request could not be safely deduplicated; retry with a new Idempotency-Key. |
| 502/504 | upstream_error / upstream_timeout | The video provider is unavailable or did not accept the request before timeout. |
| 503 | video_unavailable | Router video generation is not configured or is temporarily unavailable. |
Pricing & credits
Usage is billed in credits, where 1000 credits equal 1 USD of deposit.
A request reserves worst-case credits before it runs, settles against the usage the model actually reported, and refunds the difference. Failed requests are not billed — credits settle only on a successful response. Current per-model rates are on the Models page.
Listing models
The same catalog the Models page shows, as JSON.
curl https://router.qtum.ai/v1/models -H "Authorization: Bearer $QTUM_API_KEY"Error codes
Errors return { "error": { "code", "message" } }. The HTTP status and the stable code string are below. The Anthropic endpoint returns Anthropic's native { "type": "error", "error": { "type", "message" } } envelope instead.
| HTTP | Code | Meaning |
|---|---|---|
| 400 | invalid_params | Malformed body, or a missing / invalid model field |
| 400 | invalid_model | Model id is not available for that endpoint |
| 401 | unauthorized | Missing, invalid, or revoked API key |
| 402 | insufficient_balance | Insufficient credits — top up to continue |
| 404 | not_found | Video task id not found (video status / content polling) |
| 500 | internal_error | Internal error (e.g. pricing misconfigured) — not billed |
| 502 | upstream_error | The gateway could not reach the model provider — not billed |
| 503 | video_unavailable | Video generation is temporarily unavailable |
| 504 | upstream_timeout | The gateway timed out reaching the provider (e.g. an input image too large to upload in time) — not billed |
Ready to call it
Create a key in the console, then point your client at https://router.qtum.ai/v1.