A flat-rate 'unlimited tokens' inference API running open Llama-family models on its own GPUs — no token metering at all, with per-day request caps instead.
Pricing: Free (Lite), $5 (Core), $10 (Plus), $20 (Pro), $80/mo (Max — unlimited requests). All tiers: unlimited tokens per request; daily request caps and RPM limits below Max.
Pick Awan LLM if you want the cheapest truly-unmetered tokens on open Llama models — for bulk text or hobby projects, $5–20/month is unbeatable. Pick Standard Compute if you're running a real coding agent: it needs tool calling, frontier models and long context.
Standard Compute is an OpenAI-compatible API with frontier-model compute at a flat monthly price (from $39/mo) — no per-token billing, no rate-limit windows. Under extreme sustained load requests are paced smoothly instead of erroring or charging more.
Standard Compute is OpenAI-compatible, so any tool or SDK that lets you set a custom base URL migrates in minutes:
Base URL = https://api.stdcmpt.com/v1 API key = your Standard Compute key Model = standardcompute
Setup guides for every major agent — OpenClaw, Hermes, OpenCode, Cursor, Cline, Aider and more — on the integrations page. Free tier to test it, no card required.
Awan is flat-rate on open Llama models with daily request caps; Standard Compute is flat-rate across the full frontier (GPT-5.6, Claude, Gemini, DeepSeek, Qwen) with smart routing and no request-count caps. They solve different problems: cheapest possible tokens versus agent-grade capability at a fixed price.
Technically the endpoint is OpenAI-compatible, but tool/function calling isn't documented in Awan's API — and coding agents lean on it constantly. Test before committing; for tool-heavy agents a provider with documented tool support is the safer base.
The flat-rate alternative to Awan LLM