Z.ai · updated 2026-08-29

Run GLM 5.2: providers, pricing and flat-rate access

Z.ai's mid-2026 coding and long-horizon-agent model — the generation before GLM-5.3, still widely run for its strong coding quality at mid-tier prices.

GLM 5.2 runs cheapest via Z.ai's GLM Coding Plan (from a few dollars a month promo, up to ~$160/mo — but 5.2 burns 2–3x quota vs routine models), per-token via Z.ai's API ($1.40/$4.40 per M tokens) or OpenRouter, or flat-rate via Standard Compute, where GLM 5.2 is one of the default workhorse models in the routed pool.

Your options, honestly

GLM Coding Plan (Z.ai)~$3–160/moCheapest single-vendor entry — but GLM-only, prompt-count windows, and GLM-5.2 consumes 3x quota at peak / 2x off-peak, so plans exhaust faster than the headline suggests.
Z.ai API / OpenRouterPay-per-token$1.40/$4.40 per M tokens first-party (cached input ~$0.26); full control, bill scales with agentic context growth.
Standard Compute$39+/mo flatGLM 5.2 is a default routed model in the pool — it carries a large share of routine steps with frontier models behind it. One price, no quota multipliers, no pinning.
the flat-rate trade-off

If GLM alone covers your work, Z.ai's own plan is honestly the cheapest entry — just read the quota multipliers before trusting the headline price. The flat-rate case is mixed workloads: GLM 5.2 for routine steps and frontier models for hard ones, without managing two subscriptions or a quota calendar.

Running it in your agent

OpenCode, Cline, Aider, Roo and Continue all run GLM 5.2 through any OpenAI-compatible endpoint — set the base URL (api.stdcmpt.com/v1 for flat-rate) and go. On Standard Compute it's in the default routed pool, so 'standardcompute' reaches it with zero config.

FAQ

GLM 5.2 or GLM 5.3?
5.3 is the newer release and leads on hard agentic coding; 5.2 stays close on routine work at slightly lower per-token rates. If you don't want to manage the choice, routed plans serve both and pick per-request.
Why does my GLM Coding Plan run out so fast on 5.2?
GLM-5.2 draws 3x quota during peak hours and 2x off-peak (temporary off-peak promotions aside), so a plan sized for routine models depletes in a third of the time. Flat-rate multi-model plans don't window usage this way.

More: how flat-rate LLM APIs work · per-agent guides for OpenCode, Claude Code and Cline · model comparisons