A flat-rate unlimited-token subscription for open-weight models (DeepSeek, GLM, Kimi, Qwen and thousands more), with agent-runtime sandboxes on the higher tiers.
Pricing: Premium $25/mo (any open-weight model size, 4 concurrent connections, 32K max context), Agent Standard $100 (8 connections, 256K context, agent runtime), Agent Pro $200. Unlimited tokens on all tiers. (Verified 2026-07-17.)
Featherless and Standard Compute are the two honest 'unlimited tokens' options, split by model philosophy: Featherless gives you the open-weight universe with pinning, concurrency and context caps; Standard Compute gives you a frontier-model pool with smart routing and speed-based tiers. Pick by which models your agent actually needs.
Standard Compute is an OpenAI-compatible API with frontier-model compute with no usage limits at a flat monthly price (from $39/mo) — no per-token billing, no rate-limit windows. Under extreme sustained load requests are paced smoothly instead of erroring or charging more.
Standard Compute is OpenAI-compatible, so any tool or SDK that lets you set a custom base URL migrates in minutes:
Base URL = https://api.stdcmpt.com/v1 API key = your Standard Compute key Model = standardcompute
Setup guides for every major agent — OpenClaw, Hermes, OpenCode, Cursor, Cline, Aider and more — on the integrations page. Free tier to test it, no card required.
If open-weight quality (DeepSeek/GLM/Kimi-class) satisfies your tasks and your contexts fit the caps, Featherless is honest unlimited value. If your agent leans on frontier-model quality or big context, Standard Compute serves that class of model without a token meter.
Both drop per-token billing entirely. Featherless bounds usage with concurrency (4-8 connections) and context caps; Standard Compute has no such caps — full frontier-model access, and under extreme sustained load it paces smoothly rather than rejecting. The choice is which constraint fits: open-weight-only with concurrency/context limits, or full-frontier with fair-use pacing.
The flat-rate, no-usage-limit alternative to Featherless