
Z.ai Opens GLM-5.3 for Coding and Agent Work
On August 14, Beijing-based Z.ai released GLM-5.3, a model built for coding and agent work. Per-token pricing for it has not been published: the official price table still stops at GLM-5.2.
Why does the previous version's rate have to stand in for now
The base is unchanged: roughly the same 700 billion parameters as June's GLM-5.2. The gains came from post-training after the base stage.
The reference point for a budget is the current GLM-5.2 rate: $1.40 per million input tokens and $4.40 per million output tokens. That is the price of the previous model, not the new one. A budget built on it works as an upper or lower bound and gets refined once the rate is published.
How should a coding budget be calculated
Open weights add a second delivery option: the model can be deployed on the company's own infrastructure. Data then never leaves the company perimeter. Both scenarios are worth costing out — hardware rental against the API rate.
A budget calculator helps estimate the monthly bill at real volumes. Rates are entered manually; volume comes from last month's logs.
What is still unknown
Per-token pricing for GLM-5.3 has not been announced. The commercial licensing terms have not been stated either, nor have the hardware requirements for running the model locally. Also open: paying for the API from Russia, and whether Russian cloud providers offer the model.
Let’s discuss your project?
Tell us about your process — we’ll suggest where AI pays off fastest.
Related articles

OpenAI Ships GPT-6 Astra: 63% Lower Cost Per Task
OpenAI unveiled GPT-6 Astra: a 12-point lead on the domain-specific benchmark and a claimed 63% lower price per task. Agent budgets are being recalculated.

Sber launches GigaAgent: an agent pilot inside the Russian perimeter
Sber has opened access to an autonomous general-purpose AI agent. The launch runs through Cloud.ru Agents Space, with a RUB 4,000 starter grant for new users.