Alibaba · updated 2026-08-29
Alibaba's mid-2026 Qwen generation (3.8-Max and 27B variants) — top-tier open-weight coding performance with aggressive per-token pricing.
Qwen 3.8 runs cheapest per-token through providers like DeepInfra or Alibaba's own API (~$0.4/M input class), most flexibly through OpenRouter, and most predictably through a flat-rate plan like Standard Compute, where Qwen 3.8 serves routine agent steps inside the routed pool at a fixed monthly price.
Qwen 3.8's per-token price is genuinely low, so light users should just pay per token. The flat-rate case starts when your agent runs hours daily — context growth makes even cheap tokens add up, and a fixed price removes the variance.
All OpenAI-compatible agents (OpenCode, Cline, Aider, Continue, Roo) run Qwen 3.8 with a base-URL swap. On Standard Compute, model 'standardcompute' routes to it automatically for the steps where it's the right tool.
More: how flat-rate LLM APIs work · per-agent guides for OpenCode, Claude Code and Cline · model comparisons