Router4 min read

Point Claude Code at the Qtum Router

Two environment variables put Claude Code on the Qtum AI Router: one prepaid balance, every Claude model, and a per-request record of what a session spent. With the auth-header gotcha that causes most of the 401s.

Coding agents have an awkward billing shape. A session is not one request — it is dozens, each one resending a context that grows as the agent reads more of your codebase. The token count is unpredictable before you start and invisible while it runs, and the bill arrives attached to whichever account you happened to configure.

The Qtum AI Router speaks the Anthropic Messages API, so Claude Code and the tools around it can point at it directly: one prepaid balance, one usage page, and a per-request record of what each session actually spent.

Setup

Two environment variables:

export ANTHROPIC_BASE_URL=https://router.qtum.ai
export ANTHROPIC_AUTH_TOKEN=$QTUM_API_KEY

That is the whole configuration. The same works for claude-code-router and anything else that reads those variables.

The one thing that trips people up

Note ANTHROPIC_AUTH_TOKEN, not ANTHROPIC_API_KEY. The two are not interchangeable — they select different auth schemes:

  • ANTHROPIC_AUTH_TOKEN sends Authorization: Bearer …, which is what Claude Code and claude-code-router use.
  • ANTHROPIC_API_KEY sends x-api-key: …, which is what the anthropic Python and JavaScript SDKs use.

The Router accepts both. Set the one your tool reads, and if you get a 401 with a key you know is good, it is almost always this.

Note also that the base URL for this surface stops at the host, with no /v1 — the SDK appends /v1/messages itself. (The OpenAI-compatible surface on the same Router does want /v1 on the end. They differ because the SDKs differ.)

Which model to point it at

Every Claude model in the catalog, with today's prices per million tokens. The context column matters for agents more than for anything else, because a long session is mostly input:

ModelInputOutputContextMax output
claude-haiku-5-5-ab$0.02$0.121M128K
claude-sonnet-4-6-ab$0.53$2.651M64K
claude-opus-5-5-kiro-ab$0.64$3.181M128K
claude-opus-5-ab$0.79$3.971M128K
claude-opus-4-7-ab$0.88$4.411M128K
claude-opus-5-5-ccmax-ab$0.95$4.751M128K
claude-fable-5-1-ab$2.38$11.88200K64K
claude-opus-5-bedrock$5.40$27.001M128K

One thing worth noticing before you pick: the same model family often appears more than once at very different prices, because the catalog carries more than one upstream route to it. claude-opus-5-ab and claude-opus-5-bedrock are both Opus 5, and one is roughly seven times the other. It costs nothing to read the live model table before you set the variable.

What a session actually costs

Take a realistic session: forty exchanges, averaging 25,000 input tokens each as the context accumulates, and 1,500 output tokens each. That is 1,000,000 input and 60,000 output tokens — and because the arithmetic is public, you can price it in advance rather than discovering it:

ModelInputOutputSession
claude-haiku-5-5-ab24 credits7 credits≈ $0.03
claude-sonnet-4-6-ab529 credits159 credits≈ $0.69
claude-opus-5-ab794 credits238 credits≈ $1.03
claude-opus-5-bedrock5,400 credits1,620 credits≈ $7.02

Input dominates, which is the thing to internalise about agents. The output of a coding session is a few thousand tokens of diff; the input is your repository, sent again on every turn. That is why the input column is the one to optimise, and why a cheap model with a large context can be the right choice for the exploratory half of a task even when you want a strong model for the hard part.

How the billing protects you

Forwarding is post-paid by nature — nobody knows the token count until the response exists. The Router handles that by reserving a worst-case amount up front and correcting it afterwards:

  1. Reserve. Before the call, it deducts an estimate of the input plus max_tokens priced at the output rate. The Anthropic API requires max_tokens, so the output side is exactly bounded — the reservation can never be an underestimate.
  2. Settle. When the response completes, it charges the tokens actually used.
  3. Refund. The difference goes straight back to your balance.

Two consequences worth knowing. You cannot overdraft: a call that your balance does not cover is refused before it is forwarded, rather than running up a debt. And concurrent requests are safe — each in-flight call holds its own reservation against a row-locked balance, so an agent running several tool calls at once cannot spend the same credits twice.

Every charge lands in your usage history with the model and the request attached, so "what did yesterday's refactor cost" is a question with an answer.

Not just Claude Code

The same key answers the OpenAI-compatible surface on the same host, so any tool that reads OPENAI_BASE_URL works too:

export OPENAI_BASE_URL=https://router.qtum.ai/v1
export OPENAI_API_KEY=$QTUM_API_KEY

One balance, one usage page, both protocol families — and 91 models behind them, not only the Claude line. The integration guide covers the rest of the surface: images, speech, transcription and video.

Getting a key

Keys are created from the Router page. Top up, export two variables, and start a session — if it is a Haiku session, the first one costs about three cents.

One API key for every model

The Qtum AI Router puts text, image, audio and video models behind a single OpenAI- and Anthropic-compatible endpoint, billed in one place.