A flat-rate subscription for eight open-weight models — Kimi-K3, GLM-5.2, GLM-5.3-Flash, GLM-4.7-Flash, Nemotron-3-Super-120B, gpt-oss-120b, Qwen3.8 and an embedding model — metered by requests instead of tokens.
Pricing: $1/day or $30/month; pay-as-you-go and enterprise options exist. 500 requests per 5 hours, 1 concurrent request per model (extra concurrency packs sold separately). No token metering at all.
Synthetic is one of the most honest offers in the budget field: $30 flat, a request cap you can actually reason about, and a real privacy stance. Its constraints are structural — open-weight only, one concurrent request per model, 500 requests per 5 hours — which suits a single sequential agent and punishes parallel or high-frequency loops. Standard Compute is the alternative for exactly those: frontier-plus-open models, no request windows, no concurrency wall.
Standard Compute is an OpenAI-compatible API with frontier-model compute at a flat monthly price (from $39/mo) — no per-token billing, no rate-limit windows. Under extreme sustained load requests are paced smoothly instead of erroring or charging more.
Standard Compute is OpenAI-compatible, so any tool or SDK that lets you set a custom base URL migrates in minutes:
Base URL = https://api.stdcmpt.com/v1 API key = your Standard Compute key Model = standardcompute
Setup guides for every major agent — OpenClaw, Hermes, OpenCode, Cursor, Cline, Aider and more — on the integrations page. Free tier to test it, no card required.
Tokens are unmetered, but requests aren't: 500 per 5 hours with 1 concurrent request per model. For chat-style use that's generous; for agent loops firing tool calls every few seconds, the request budget and the concurrency cap both bind. Standard Compute meters neither.
If open-weight quality covers you, privacy matters, and you run one agent sequentially inside OpenAI-compatible tools like Roo, Cline, or Octofriend, Synthetic's $30 is excellent value. If your agent runs parallel streams, needs frontier quality, or burns more than 500 requests in 5 hours, that's the regime flat-rate without windows exists for.
The flat-rate alternative to Synthetic