Google · updated 2026-09-05
Google’s latest Flash model for fast agent workflows, with a one-million-token context window.
Run Gemini 3.8 Flash through the Google API, through OpenRouter, or with Standard Compute’s monthly compute plans. On Standard Compute, select it in the dashboard and keep your agent configured with the standardcompute model. Your monthly budget covers usage; a fixed plan price does not mean unlimited tokens.
Model documentation: Google: model documentation
A direct API can be a better fit for light or occasional usage. A monthly compute plan is useful when you want a predictable bill across several models. Premium models consume the budget faster; compare the cost of completed tasks, including retries and cached input.
In an OpenAI-compatible agent, set the base URL to https://api.stdcmpt.com/v1, paste your API key, and use standardcompute as the model. Select Gemini 3.8 Flash in Dashboard → Models. Select only this model to pin requests to it, or add up to four others for smart routing.
More: how flat-rate LLM APIs work · per-agent guides for OpenCode, Claude Code and Cline · model comparisons