Coding agents have an awkward billing shape. A session is not one request — it is dozens, each one resending a context that grows as the agent reads more of your codebase. The token count is unpredictable before you start and invisible while it runs, and the bill arrives attached to whichever account you happened to configure.
The Qtum AI Router speaks the Anthropic Messages API, so Claude Code and the tools around it can point at it directly: one prepaid balance, one usage page, and a per-request record of what each session actually spent.
Setup
Two environment variables:
export ANTHROPIC_BASE_URL=https://router.qtum.ai
export ANTHROPIC_AUTH_TOKEN=$QTUM_API_KEY
That is the whole configuration. The same works for claude-code-router and anything else that reads those variables.
The one thing that trips people up
Note ANTHROPIC_AUTH_TOKEN, not ANTHROPIC_API_KEY. The two are not
interchangeable — they select different auth schemes:
ANTHROPIC_AUTH_TOKENsendsAuthorization: Bearer …, which is what Claude Code and claude-code-router use.ANTHROPIC_API_KEYsendsx-api-key: …, which is what theanthropicPython and JavaScript SDKs use.
The Router accepts both. Set the one your tool reads, and if you get a 401 with a key you know is good, it is almost always this.
Note also that the base URL for this surface stops at the host, with no /v1 —
the SDK appends /v1/messages itself. (The OpenAI-compatible surface on the
same Router does want /v1 on the end. They differ because the SDKs differ.)
Which model to point it at
Every Claude model in the catalog, with today's prices per million tokens. The context column matters for agents more than for anything else, because a long session is mostly input:
| Model | Input | Output | Context | Max output |
|---|---|---|---|---|
claude-haiku-5-5-ab | $0.02 | $0.12 | 1M | 128K |
claude-sonnet-4-6-ab | $0.53 | $2.65 | 1M | 64K |
claude-opus-5-5-kiro-ab | $0.64 | $3.18 | 1M | 128K |
claude-opus-5-ab | $0.79 | $3.97 | 1M | 128K |
claude-opus-4-7-ab | $0.88 | $4.41 | 1M | 128K |
claude-opus-5-5-ccmax-ab | $0.95 | $4.75 | 1M | 128K |
claude-fable-5-1-ab | $2.38 | $11.88 | 200K | 64K |
claude-opus-5-bedrock | $5.40 | $27.00 | 1M | 128K |
One thing worth noticing before you pick: the same model family often appears
more than once at very different prices, because the catalog carries more than
one upstream route to it. claude-opus-5-ab and claude-opus-5-bedrock are
both Opus 5, and one is roughly seven times the other. It costs nothing to read
the live model table before you set the variable.
What a session actually costs
Take a realistic session: forty exchanges, averaging 25,000 input tokens each as the context accumulates, and 1,500 output tokens each. That is 1,000,000 input and 60,000 output tokens — and because the arithmetic is public, you can price it in advance rather than discovering it:
| Model | Input | Output | Session |
|---|---|---|---|
claude-haiku-5-5-ab | 24 credits | 7 credits | ≈ $0.03 |
claude-sonnet-4-6-ab | 529 credits | 159 credits | ≈ $0.69 |
claude-opus-5-ab | 794 credits | 238 credits | ≈ $1.03 |
claude-opus-5-bedrock | 5,400 credits | 1,620 credits | ≈ $7.02 |
Input dominates, which is the thing to internalise about agents. The output of a coding session is a few thousand tokens of diff; the input is your repository, sent again on every turn. That is why the input column is the one to optimise, and why a cheap model with a large context can be the right choice for the exploratory half of a task even when you want a strong model for the hard part.
How the billing protects you
Forwarding is post-paid by nature — nobody knows the token count until the response exists. The Router handles that by reserving a worst-case amount up front and correcting it afterwards:
- Reserve. Before the call, it deducts an estimate of the input plus
max_tokenspriced at the output rate. The Anthropic API requiresmax_tokens, so the output side is exactly bounded — the reservation can never be an underestimate. - Settle. When the response completes, it charges the tokens actually used.
- Refund. The difference goes straight back to your balance.
Two consequences worth knowing. You cannot overdraft: a call that your balance does not cover is refused before it is forwarded, rather than running up a debt. And concurrent requests are safe — each in-flight call holds its own reservation against a row-locked balance, so an agent running several tool calls at once cannot spend the same credits twice.
Every charge lands in your usage history with the model and the request attached, so "what did yesterday's refactor cost" is a question with an answer.
Not just Claude Code
The same key answers the OpenAI-compatible surface on the same host, so any tool
that reads OPENAI_BASE_URL works too:
export OPENAI_BASE_URL=https://router.qtum.ai/v1
export OPENAI_API_KEY=$QTUM_API_KEY
One balance, one usage page, both protocol families — and 91 models behind them, not only the Claude line. The integration guide covers the rest of the surface: images, speech, transcription and video.
Getting a key
Keys are created from the Router page. Top up, export two variables, and start a session — if it is a Haiku session, the first one costs about three cents.
