Alibaba · updated 2026-08-29
Alibaba's fast multimodal Qwen tier — text, image and video understanding with a 1M-token context at some of the lowest per-token rates on the market ($0.03/$0.13 per M tokens for short-context requests).
Qwen 3.7 Flash is nearly free per-token for light use — $0.03/M input on Alibaba's international endpoint (short-context tier; longer prompts bill higher) — so light users should simply use Alibaba's API or OpenRouter. Standard Compute includes it in the flat-rate routed pool as a high-speed workhorse for routine and multimodal steps.
Honestly: at $0.03/M input, no one needs a subscription for a Flash-only workload — pay per token. The flat-rate case is what surrounds it: agents that use Flash for volume but need frontier quality on hard steps, where the frontier tokens — not the Flash ones — are what make per-token bills unpredictable.
All OpenAI-compatible agents run Qwen 3.7 Flash with a base-URL swap. On Standard Compute, 'standardcompute' routes fast routine steps (including image/video understanding) to it automatically.
More: how flat-rate LLM APIs work · per-agent guides for OpenCode, Claude Code and Cline · model comparisons