Z.ai · updated 2026-08-29
Z.ai's mid-2026 coding and long-horizon-agent model — the generation before GLM-5.3, still widely run for its strong coding quality at mid-tier prices.
GLM 5.2 runs cheapest via Z.ai's GLM Coding Plan (from a few dollars a month promo, up to ~$160/mo — but 5.2 burns 2–3x quota vs routine models), per-token via Z.ai's API ($1.40/$4.40 per M tokens) or OpenRouter, or flat-rate via Standard Compute, where GLM 5.2 is one of the default workhorse models in the routed pool.
If GLM alone covers your work, Z.ai's own plan is honestly the cheapest entry — just read the quota multipliers before trusting the headline price. The flat-rate case is mixed workloads: GLM 5.2 for routine steps and frontier models for hard ones, without managing two subscriptions or a quota calendar.
OpenCode, Cline, Aider, Roo and Continue all run GLM 5.2 through any OpenAI-compatible endpoint — set the base URL (api.stdcmpt.com/v1 for flat-rate) and go. On Standard Compute it's in the default routed pool, so 'standardcompute' reaches it with zero config.
More: how flat-rate LLM APIs work · per-agent guides for OpenCode, Claude Code and Cline · model comparisons