
Three cents per task: where DeepSeek V4-Flash cuts costs
Running a single routine task on DeepSeek V4-Flash was measured at 3 cents. Claude Fable 5 came out at $3.15 on the same benchmark. The gap is roughly 105-fold. Numbers like these widen the list of jobs worth handing to a model.
The figures below are indicative estimates for a company of the relevant profile, not a report on a specific deployment. Each one comes with its method, so the math can be repeated on internal data.
Why does the price list show only half the picture
The published pricing is simple: $0.14 per million input tokens and $0.28 per million output tokens. On that line alone, the model ranks among the three cheapest on the market.
Artificial Analysis counts differently. Its analysts measured the cost of a complete task in advanced reasoning mode at maximum compute load. The build tested was V4-Flash-0731.
The two methods diverge sharply. A model with low list pricing can burn more tokens while searching for a solution. That is why the gap on the price list and the gap in cost per task do not match.
Can ten thousand emails a month really fit into two thousand rubles
Dollars per million tokens does not answer the question "what does it cost." Hence the math.
The assumptions. Request profile: 10,000 input tokens and 2,000 output tokens. That is an email with its thread history plus a short structured reply. Pricing is the published $0.14 and $0.28. One such request works out to roughly 0.2 cents.
The volume. 10,000 requests a month — about 500 queries per working day.
The result. Around $20 a month. At 80 RUB to the dollar, roughly RUB 1,600. Caching of repeated prompt segments is not factored in; with it the bill drops further.
A second calculation covers harder work. A thousand tasks in advanced reasoning mode, per the Artificial Analysis benchmark, costs $30, about RUB 2,400. The same thousand on Claude Fable 5, on the same benchmark, costs $3,150, about RUB 252,000.
What the math leaves out: prompt development, testing, acceptance, manual proofreading of responses. Those line items usually drive the budget, while the bill for requests turns out to be the smaller share. The AI budget calculator helps sketch the full picture.
Where are 50 points enough
On the Artificial Analysis Intelligence Index the model scored 50 out of 100. Google Gemini 3.6 Flash scores the same. Kimi K3 sits higher at 57, as do the OpenAI and Anthropic models.
The score on its own says nothing. Here is what it does say: on ticket classification, field extraction from an invoice, or a short reply drawn from a manual, the difference between 50 and 57 points usually disappears into the noise. Answers like those are checked by eye in a second.
Long reasoning chains are a different story. A legal conclusion, a contract-terms review, a multi-step analysis with branching — there every extra point converts into a number of corrections. The developer is betting on the price-to-quality ratio and preparing the stronger V4-Pro separately.
Which profiles does this change the math for
High-volume routine processing. Source documents, inbound tickets, extracting item codes from emails. At this pricing, five-figure monthly volume stops being constrained by the API bill. This week's task: count the real number of requests and the average text length in each.
Shelved pilots. Some projects stalled at the "request costs" line in the budget. At three cents per task, that line stops being an argument against. What to do: pull out the old estimate and rerun it at the new pricing.
Agent loops. Where a single result costs dozens of model calls, the price per request is multiplied by the depth of the chain. A cheap step makes long scenarios calculable. What to do: measure how many calls go into one finished result.
A similar shift already played out during OpenAI's API price cuts. Price competition has become a market storyline in its own right: what gets counted now is not generation quality but the cost per million requests.
Why do savings only appear at volume
There is a flip side. At a thousand requests a month, the bill on the current model already runs to hundreds of rubles. Migration — rewritten prompts, new tests, another round of acceptance — eats the gain in the first month alone. The threshold below which switching models pays off poorly is a calculation each company runs on its own figures: monthly savings against the one-time cost of migrating.
The second condition is a task with a verifiable result. Where an answer cannot be checked in a second, a cheap call makes no difference: the time goes into proofreading.
The third is order in the data. If documents sit across three systems and an inbox, the price per request has no bearing on the outcome at all.
What does the source say about Russia
Payment terms from Russia are not covered in the release announcement. Neither are access routes, data residency, or compliance with personal-data requirements. There is nothing to speculate on here: it gets verified on the access provider's page before the first payment.
A separate question is running the model on-premise. License terms and hardware requirements are not stated in the source.
What remains unknown: the timeline for a full V4-Pro release; license terms for the weights; the context limit and response speed of V4-Flash; pricing for cached input from the official price list; how the model handles Russian-language documents — the source contains no public Russian-language benchmarks.
The model did not become smarter than the leaders of the ranking. It made cheap the tasks that never needed a powerful mind in the first place. For high-volume work, that is usually enough.
Let’s discuss your project?
Tell us about your process — we’ll suggest where AI pays off fastest.
Related articles

Qwen3.8-Max: a million tokens of context at $2 per input
Alibaba has released a flagship model with 2.4 trillion parameters and a context window of up to 1 million tokens. What that changes in the budget for long documents.

The Real Cost of AI Adoption: Nine Budget Lines and Three Hidden Ones
The range from RUB 500,000 to several million comes down to nine cost lines. Here the budget is taken apart piece by piece, including the three lines that rarely appear in a vendor proposal.