Gonka sets the lowest price for DeepSeek V4 Flash. How long will this window last?

Price Per Token, an independent API tracker, now ranks Gonka as the cheapest of 23 providers offering DeepSeek V4 Flash. At roughly $0.00015 per million tokens, Gonka is hundreds of times cheaper than the next provider on the list — and no, that is not a typo. But this is not yet a mature market price: dynamic pricing has not reached devshards. For early users, that creates a rare opportunity — while it lasts.
The gap is so large that the ranking alone does not tell the full story. Before getting to why it exists — and why it may narrow — we need to look at what the tracker is actually comparing.
The number on the board
Price Per Token places Gonka in the same table as 22 other API providers offering the non-reasoning version of DeepSeek V4 Flash.
The first thing you notice is the $0.000 shown in both price columns.

Gonka at the top of the DeepSeek V4 Flash provider comparison, sorted by input price. Source: Price Per Token, accessed August 17, 2026.
That does not mean inference is literally free. The table rounds prices to three decimal places, while the underlying Gonka price is roughly $0.00015 per million tokens. Even without that rounding, the gap remains extraordinary: the next listed offers begin at $0.061 per million input tokens and $0.123 per million output tokens — more than 400 times higher for input and about 800 times higher for output.
The model page also lists a one-million-token context window, tools, and caching. In other words, the price is relevant not only for quick experiments, but also for applications working with large contexts and repeated requests.
What the ranking establishes is simple: Gonka is currently listed far below every other provider on the page. What it does not explain is how that price became possible — or whether it represents an equilibrium that can last.
Before dynamic pricing kicks in
Gonka is not a conventional API company that buys computing capacity and resells it at a fixed margin. It is a decentralized inference network in which independent hosts contribute GPU capacity and users create demand by sending requests.

Live inference activity across the Gonka network. Metrics shown are a point-in-time snapshot. Source: Gonka Inference Map, accessed August 17, 2026.
Dynamic pricing was part of the model from the beginning: instead of one operator choosing the price, the cost of inference should move with available capacity and demand.
Devshards — the part of Gonka’s infrastructure through which models serve inference requests — were introduced before dynamic pricing was extended to them. As a result, the user-facing price does not yet change automatically with demand and available capacity. That is why the price can stay close to zero even as demand grows.
For early users, the upside is obvious: workloads that would normally require a meaningful API budget can be tested for almost nothing. The trade-off is that price cannot yet help regulate demand.
During periods of particularly high activity, users may encounter congestion and less predictable access. For those testing or prototyping workloads, that may be a reasonable compromise at the current price.

At the time of capture, DeepSeek V4 Flash was being served by 11 hosts. Source: https://tracker.gonka.vip/stats, accessed August 17, 2026.
Dynamic pricing is expected to reach devshards in the coming months. When it does, the price will likely rise and start responding to demand and available capacity. Until then, hosts can also vote to raise the price manually if demand begins to outpace the network’s ability to serve it. That means the current rate could change even before dynamic pricing arrives.
For now, early users still have an unusually cheap window to experiment. But that raises an obvious question: if users are paying almost nothing, how is the cost of that compute being covered?
After the window closes
API fees are only one part of the network’s economics today. Hosts invest in capacity ahead of mature demand and receive network rewards for contributing it. This allows Gonka to build supply first, while the price paid by users can remain below the full cost of providing the infrastructure.
That gap is not meant to last forever. As demand grows and dynamic pricing reaches devshards, more of the cost of inference will probably be reflected in the API price. Hosts can also vote to raise the price earlier if demand begins to outpace capacity. Today’s near-zero rate is therefore better understood as an early-market window than as a permanent benchmark.
But a higher price does not necessarily mean an expensive one. Gonka can draw capacity from independent hosts in different regions, without a single API company setting a fixed margin for the entire network. Competition between suppliers could help decentralized inference remain cheaper than many centralized alternatives after the current gap narrows. But that is an advantage, not a guarantee.
For builders, using Gonka does not have to be an all-or-nothing decision. DeepSeek V4 Flash is also available through OpenRouter, so applications can send cost-sensitive traffic through Gonka and fall back to the same model elsewhere when availability or response time matters more than price.
The current window is especially useful for:
- model evaluations and prompt testing;
- long-context document processing;
- batch jobs and backfills;
- agent and tool prototypes.
Price is still only one part of the evaluation. Latency, reliability, and performance should be tested separately, especially for production workloads. The real question is not whether today’s price lasts forever, but how much of Gonka’s advantage remains once the market begins pricing demand in real time.
Access to Gonka is provided through independent community brokers using OpenAI-compatible endpoints. The current list and setup instructions are available in the developer documentation.
Your support helps keep the blog independent
It helps us spend more time on analysis and original stories about the network.