Roo Code's bill multiplier is its mode system: an Architect pass, a Code pass, and a Debug loop are three separate model conversations for one piece of work, each re-sending context. Putting cheap models on the low-stakes modes and reserving the strong model for Code mode cuts more than any single other change.
Roo Code lets you configure a different provider/model per mode. Architect and Ask run fine on a budget model; give Code mode the strong model. This alone routinely halves the bill on plan-heavy workflows.
Long exploratory planning chats are the hidden cost center. Front-load the constraints in your first message, get the plan, switch modes — every extra Architect round-trip re-bills your project context.
Point Roo Code at the directories that matter instead of letting it roam the whole monorepo. Smaller working sets mean smaller re-sent context on every mode's every message.
Three passes per task is a structural multiplier that per-token billing punishes and capped multipliers exhaust. A plan with no usage limits turns the multiplier into a non-event.
The provider-agnostic tactics (prompt caching, retry budgets, batch APIs) are in the general playbook.
Usually the mode system: if you plan in Architect, implement in Code, and iterate in Debug, you've run roughly three conversations where a single-mode agent ran one. Per-mode model assignment is the built-in fix.
Code mode — it's where output quality directly becomes your diff. Architect produces plans (cheap models plan adequately with a good prompt), and Debug is iterative enough that a mid-tier model usually suffices.
The structural version of all of this: run Roo Code on a flat monthly price with unlimited tokens, and the bill stops being a variable to manage. 2-minute Roo Code setup → · Best models for Roo Code →