Alibaba · updated 2026-09-05

Run Qwen 3.7 Flash: providers, pricing and flat-rate access

Alibaba's fast multimodal Qwen tier — text, image and video understanding with a 1M-token context at some of the lowest per-token rates on the market ($0.03/$0.13 per M tokens for short-context requests).

Qwen 3.7 Flash is nearly free per-token for light use — $0.03/M input on Alibaba's international endpoint (short-context tier; longer prompts bill higher) — so light users should simply use Alibaba's API or OpenRouter. Standard Compute includes it in the flat-rate routed pool as a high-speed workhorse for routine and multimodal steps.

Your options, honestly

Alibaba Model Studio APIPay-per-tokenFirst-party and extremely cheap ($0.03/$0.13 per M tokens under 32K context; tiered pricing above that); you manage keys, tiers and spend alerts.
OpenRouterPay-per-tokenSame model with unified billing and fallback; a small fee on credits, less tier bookkeeping.
Standard Compute$19+/mo flat3.7 Flash in the routed pool for fast routine and vision/video steps, with frontier models behind it — one price. Choose your model selection in the dashboard.
the flat-rate trade-off

Honestly: at $0.03/M input, no one needs a subscription for a Flash-only workload — pay per token. The flat-rate case is what surrounds it: agents that use Flash for volume but need frontier quality on hard steps, where the frontier tokens — not the Flash ones — are what make per-token bills unpredictable.

Running it in your agent

All OpenAI-compatible agents run Qwen 3.7 Flash with a base-URL swap. On Standard Compute, add Qwen 3.7 Flash to your model selection and keep 'standardcompute' set in your agent.

FAQ

Is Qwen 3.7 Flash actually good, or just cheap?
It's a genuine vision-language reasoning model with a 1M context — strong for routine coding steps, extraction and multimodal work. It won't match frontier models on hard reasoning, which is exactly the split a routed pool manages per-request.
Why did my Qwen Flash bill exceed the headline rate?
The $0.03/$0.13 rate applies to requests under 32K tokens; longer contexts bill on higher tiers, and agentic loops grow context every step. Still cheap in absolute terms — but the headline number isn't the whole schedule.

More: how flat-rate LLM APIs work · per-agent guides for OpenCode, Claude Code and Cline · model comparisons