[UTC+3]

Z.ai Opens GLM-5.3 for Coding and Agent Work

August 17, 2026 · 1 minNewsReleases

On August 14, Beijing-based Z.ai released GLM-5.3, a model built for coding and agent work. Per-token pricing for it has not been published: the official price table still stops at GLM-5.2.

Why does the previous version's rate have to stand in for now

The base is unchanged: roughly the same 700 billion parameters as June's GLM-5.2. The gains came from post-training after the base stage.

The reference point for a budget is the current GLM-5.2 rate: $1.40 per million input tokens and $4.40 per million output tokens. That is the price of the previous model, not the new one. A budget built on it works as an upper or lower bound and gets refined once the rate is published.

$1.40
GLM-5.2: per million input tokens
$4.40
GLM-5.2: per million output tokens
700 billion
parameters, the same count as GLM-5.2
Source: Z.ai official price table as of August 17, 2026; the GLM-5.3 rate is not published

How should a coding budget be calculated

Open weights add a second delivery option: the model can be deployed on the company's own infrastructure. Data then never leaves the company perimeter. Both scenarios are worth costing out — hardware rental against the API rate.

A budget calculator helps estimate the monthly bill at real volumes. Rates are entered manually; volume comes from last month's logs.

What is still unknown

Per-token pricing for GLM-5.3 has not been announced. The commercial licensing terms have not been stated either, nor have the hardware requirements for running the model locally. Also open: paying for the API from Russia, and whether Russian cloud providers offer the model.

Source: incrussia.ru