Three new models went live on the Qtum AI Router this week, and the headline is a number: 20 credits per million tokens — $0.02 — and the same rate for input and output.
That is not a rounding difference against the rest of the catalogue. It is one to two orders of magnitude.
| Model | Input / 1M | Output / 1M |
|---|---|---|
| DeepSeek V4 Flash · Gonka Network | 20 | 20 |
| MiniMax M2.7 · Gonka Network | 20 | 20 |
| GLM 5.3 Flash · Gonka Network | 20 | 20 |
| DeepSeek V4.1 Flash · 官方 | 176 | 706 |
| MiniMax M3 | 360 | 1,440 |
| GLM 5.1 | 990 | 3,961 |
| GPT-5.5 | 4,200 | 25,200 |
Credits, live from the models page. 1,000 credits = USD 1.
They are cheap because of where they run.
What Gonka is
Gonka is a decentralized network for AI compute with its own token, GNK, launched in 2025. Instead of renting GPUs from a cloud provider, it coordinates machines contributed by independent operators — anything from a single card to a datacentre — and pays them for the inference they serve.
The part that makes it more than a rented-GPU marketplace is the consensus design, which the project calls Proof of Compute. In an ordinary proof-of-work chain, the GPUs burn their cycles on a puzzle that exists only to secure the chain. Gonka inverts that: hosts prove their capacity in a short synchronised window called a Sprint, and the rest of the epoch — 15,391 blocks, roughly 23 hours — is free for useful work. The proof and the product are the same hardware doing the same kind of maths.
A request is routed so that no single host both receives and answers it:
- The client reaches a randomly chosen host, acting as Transfer Agent.
- That agent randomly picks an Executor from the other active hosts and forwards the input, while recording the input on-chain in parallel with the computation.
- The Executor's ML node runs the inference and returns the output back through the agent to the client.
- The Executor records a validation artifact on-chain. Validation happens after the answer has been delivered, so verification never sits in the latency path.
An executor caught returning bad work loses its rewards for the entire epoch, and the client is refunded. That is the economic substitute for trusting a provider's SLA.
On scale, the project reports the computational equivalent of more than 6,000 NVIDIA H100s, and a conference presentation by a co-creator cited 4,000+ GPUs across 26 countries as of April 2026. Bitfury announced a $50 million commitment to the network in December 2025. Those figures come from the project and its backers rather than from independent measurement, so treat them as self-reported — but the direction is clear, and it is the reason the inference is priced the way it is.
The three models
All three are the fast, efficient tier of their family — the ones built for volume rather than for the hardest single question you have. Each is served with a 200,000-token operational context ceiling, and each supports streaming and tool calling through the Router's OpenAI-compatible endpoint.
deepseek-ai/DeepSeek-V4-Flash-0731 DeepSeek V4 Flash · Gonka Network
MiniMaxAI/MiniMax-M2.7 MiniMax M2.7 · Gonka Network
zai-org/GLM-5.3-Flash GLM 5.3 Flash · Gonka Network
DeepSeek V4 Flash is DeepSeek's throughput tier — the one most people reach for when a task is well-specified and the volume is high: classification, extraction, summarisation, straightforward code edits.
MiniMax M2.7 comes from the family MiniMax has been pointing squarely at agent work — long multi-step runs where the model calls tools, reads results and keeps going.
GLM 5.3 Flash is Zhipu's fast tier, and strong multilingual coverage is what that family is known for.
We are not going to publish our own benchmark numbers here. At $0.02 per million tokens you can run your own evaluation on your own data for less than the price of reading about someone else's.
Spotting them in the catalogue. The
· Gonka Networksuffix in the display name is how you tell which channel serves a model. It matters: the catalogue carries aDeepSeek V4 Flash · 火山方舟at 212 / 636 credits right next to the Gonka one at 20 / 20. Same model family, same name until the suffix, more than ten times the cost.
Where the saving actually lands
The flat rate is the part worth dwelling on. Every other model in the catalogue charges three to six times more for output than for input, because output is the expensive half to generate. Gonka charges the same for both.
That reshapes which workloads are cheap. Take one ordinary agent turn — 20,000 tokens of context in, 2,000 tokens out:
| Model | Cost of that one turn | vs Gonka |
|---|---|---|
| Any of the three, on Gonka | 0.44 credits | — |
| DeepSeek V4.1 Flash · 官方 | 4.94 credits | 11× |
| MiniMax M3 | 10.08 credits | 23× |
| GLM 5.1 | 27.72 credits | 63× |
| GPT-5.5 | 134.40 credits | 305× |
Per million tokens of output alone, the gap is starker still: $0.02 against $0.71 for the same DeepSeek family through a conventional channel, $3.96 for GLM 5.1, $25.20 for GPT-5.5.
So the workloads that change character are the generation-heavy ones:
- Agent loops, where every turn replays the context and emits a few hundred tokens of reasoning or a tool call. Hundreds of turns stop being a budget decision.
- Batch processing — tens of thousands of documents to classify, translate, summarise or restructure.
- Synthetic data, where output volume is the entire point.
- Evaluation and iteration: running the same prompt fifty ways to see which holds up, which at flagship prices most people simply do not do.
A good pattern is to draft on Gonka and finish on a flagship. The Router makes that one string, not one integration — see running Claude Code on the Router and the OpenClaw guide for agents that switch models with a single command.
Calling them
Nothing new to learn. Same key, same endpoint, same OpenAI-compatible wire format as every other model in the catalogue — only the model string changes:
curl https://router.qtum.ai/v1/chat/completions \
-H "Authorization: Bearer $QTUM_ROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-ai/DeepSeek-V4-Flash-0731",
"messages": [{"role": "user", "content": "Summarise this in three bullets: ..."}],
"max_tokens": 1024
}'
Streaming and tool calling work exactly as they do elsewhere on the Router:
from openai import OpenAI
client = OpenAI(base_url="https://router.qtum.ai/v1", api_key=QTUM_ROUTER_API_KEY)
stream = client.chat.completions.create(
model="zai-org/GLM-5.3-Flash",
messages=[{"role": "user", "content": "Explain Proof of Compute in two sentences."}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
Billing is settled from the token counts the upstream reports, not from an estimate, and a request that fails costs nothing. Note the ids carry the upstream organisation prefix — deepseek-ai/, MiniMaxAI/, zai-org/ — so copy them exactly from the models page.
What to know before you move a workload
Honest notes, because a cheap model you have to debug is not cheap:
- It is a shared, decentralized network. Capacity fluctuates with what the hosts are doing. Under saturation a request can come back
429; the gateway already retries transient ones with backoff, but a client that retries on429will have a better time than one that does not. For batch work, that is a non-issue. For a user-facing request with a hard latency budget, size it accordingly. - 200k is an operational ceiling, not the model's nominal context. Oversized prompts are rejected up front with
context_length_exceededrather than being silently truncated upstream. - These are efficiency-tier models. For the hardest reasoning you still want a flagship. The point is that you now have somewhere cheap to send everything that is not that.
Try it on something real
The quickest honest test: take a prompt you already run on a flagship, run it a hundred times on deepseek-ai/DeepSeek-V4-Flash-0731, and compare both the output and the bill. A hundred turns of 20k in and 2k out comes to 44 credits — four and a half cents.
Create an API key → · Models and pricing → · Integration guide →
Prices quoted are live as of publication; the models page is always authoritative. Network figures for Gonka are reported by the project and its backers.
