
DeepSeek V4.1 Flash: Agent Workloads Get Repriced
DeepSeek has released V4.1 Flash. The model uses an MoE design: 552 billion parameters in total, of which 8–16 billion switch on for any single response. That is why the price per response sits at the level of a budget tier.
What gets billed — parameters or activation?
In an MoE, only a small slice of the network handles each request. Here that slice is 8–16 billion out of 552 — roughly one thirtieth. Inference costs are calculated on the active portion, which is how a heavy model lands in a cheap tier.
Three numbers are claimed on agent benchmarks. Terminal work: 90.6 points, against 89.1 for Claude Opus 5 and 88.8 for GPT-5.6. Bug fixing in real projects: 74.2.
What does this mean for the budget?
Agent and coding workloads are the most expensive line on the bill: an AI agent makes dozens of calls to finish one task. Shifting that load onto cheaper inference moves the monthly bill further than any amount of prompt tuning.
The way to check it is straightforward: take one week of real requests, run them through both models, then compare the bill and the share of correct answers. To get a sense of the numbers in advance, use the inference cost calculator.
What is still unknown?
Token prices for Russian payers, the licence and the availability of open weights for deployment inside a company's own perimeter, the context limit, and the date the API opens.
Let’s discuss your project?
Tell us about your process — we’ll suggest where AI pays off fastest.
Related articles

OpenAI Ships GPT-6 Astra: 63% Lower Cost Per Task
OpenAI unveiled GPT-6 Astra: a 12-point lead on the domain-specific benchmark and a claimed 63% lower price per task. Agent budgets are being recalculated.

Gemini 3.7 Flash Costs Half as Much: Recalculating the Throughput Budget
Google has released Gemini 3.7 Flash — the cheap model tier has improved at code and agents, and input costs $0.75 per 1M tokens through the end of 2026.