# Agent workload report

Template version: 1.0 · September 5, 2026

This is a blank measurement worksheet, not a benchmark or a claim about any
provider. Use it with the companion [CSV](https://standardcompute.com/distribution/agent-workload-template.csv).
You can edit the files locally; neither file submits your measurements.

Prepared by Standard Compute, a commercial LLM API provider. This template is
provider-neutral: report any provider, including a local model, and include
unsuccessful runs. No purchase or account is needed to use the template.

## The question

- Decision this report supports: [e.g. which provider fits my coding workload]
- Author and affiliations: [include paid access, free credits, or sponsorship]
- Measurement dates and timezone:
- Agent and exact version:
- Provider(s), requested model IDs, and observed model IDs when available:
- Number of distinct tasks / repeated runs per task:
- Representative workload and important tasks this sample does not cover:

## Protocol

1. Choose tasks before running them. For code changes, specify the starting
   commit and a checkable acceptance condition: relevant tests, expected output,
   and any necessary human review. Passing tests is not proof of every quality
   dimension; state what they miss.
2. Keep the task, starting files, tools, instructions, and approval policy
   consistent across the providers being compared. Record model selection,
   reasoning, context, output, retry, and concurrency settings. State deliberate
   differences instead of presenting them as identical conditions.
3. Use separate fresh copies of the task's starting state. Randomize or alternate
   provider order when practical. Repeat runs to show variability. A single run
   is a case study; a few repeats still do not establish a universal ranking.
4. Record cold-cache and warm-cache runs separately. A fresh conversation does
   not prove a cold provider cache. If cache state cannot be verified, mark it
   unknown. Record cache TTL and any cache/session headers in settings notes.
5. Record all requests, retries, failed attempts, tool charges, elapsed time,
   and human interventions. Do not discard expensive failures from cost totals.
   Stop at a predeclared task time/request/spend limit and record the failure.
6. Check the agent's result against the acceptance condition. Record the actual
   evidence, rather than accepting the agent's own statement that it succeeded.
7. Capture provider usage records or billing exports. Compare those with client
   estimates. If they differ, preserve both and explain which number you use.
8. Publish only information you have permission to share. Use sanitized task
   labels and public fixtures; omit API keys, private prompts, customer code,
   email addresses, and identifiers from shared evidence.

## CSV field guide

One row is one complete task run, including its retries and subagent requests.
Use separate run IDs for separate repeats. If a run routes to several models,
write `multiple` in `observed_model`, describe the routing in `settings_notes`,
and retain a sanitized per-request breakdown when available. Do not guess the
served model from writing style or a provider's marketing page.

Leave unknown measurements blank. Use zero only when you observed zero.
Use decimal numbers without currency symbols, commas, or thousands separators.
Do not place spreadsheet formulas in imported text fields.

| Fields | Meaning |
| --- | --- |
| `schema_version` | `1.0` for this template. |
| `run_id`, `repeat_index`, `started_at_utc` | Your local run label, repeat number, and ISO 8601 UTC start. Do not use a private production identifier. |
| `agent`, `agent_version`, `provider` | Exact software and provider; local inference is a valid provider. |
| `requested_model`, `observed_model`, `api_protocol` | What you asked for, what response metadata reports, and protocol such as Chat Completions or Messages. A routed alias is not evidence of a pinned underlying model. |
| `task_id`, `task_category`, `starting_revision`, `prompt_hash` | Reproducible task identity and fixture. A hash identifies an unchanged prompt but does not make private content safe to share. |
| `cache_condition`, `settings_notes` | `cold`, `warm`, or `unknown`; reasoning/context/retry/concurrency/routing settings and deliberate differences. |
| `success`, `success_check`, `human_interventions` | `true`, `false`, or blank if unevaluated; acceptance evidence and count of manual corrections/steering actions. |
| `wall_seconds` | Full elapsed time, including retries, waits, and interventions. Record any excluded setup time in notes. |
| `requests_total`, `retries_total` | All attempted inference requests, including retries; retries are a subset of requests, not an extra amount to add. Include subagents and background requests attributable to this task. |
| `http_429_count`, `http_402_count`, `other_errors_total` | Observed request outcomes. Preserve provider error codes in notes; the same status can mean different things across providers. |
| `input_uncached_tokens`, `input_cache_read_tokens`, `input_cache_write_tokens` | Mutually exclusive normalized input categories only where the provider exposes enough detail. Some APIs include cached input in total input; others expose cache writes separately. Document the mapping. Do not subtract cache fields blindly. |
| `output_tokens`, `reasoning_tokens_reported`, `reasoning_in_output` | Provider-reported output and any reasoning detail; `reasoning_in_output` is `true`, `false`, or blank. Do not add reasoning tokens to output if already included. |
| `total_tokens_raw` | Raw provider-reported total when available. May use different conventions from normalized fields; it is not a bill by itself. |
| `currency` | ISO currency, e.g. `USD`. Do not combine currencies without documenting a conversion date and rate. |
| `provider_usage_cost` | Provider-reported metered usage amount for the run, if available. Label estimates in notes. For subscriptions this can be an accounting amount rather than cash paid. |
| `budget_debit` | Amount deducted from a subscription's included compute budget, in that budget's stated unit. Describe the unit in notes. Keep it distinct from cash and from advertised direct-API equivalents. |
| `external_tool_cost` | Search, browser, sandbox, hosting, or other directly attributable run costs not already included in `provider_usage_cost`. Explain omissions. |
| `pricing_source_url`, `prices_checked_at_utc` | Primary pricing/plan source and verification date. Record taxes, fees, credits, discounts, cache-write rates, and regional/peak rates where relevant. |
| `usage_evidence`, `notes` | Sanitized evidence reference, missing fields, uncertainty, routing mix, and anomalies. |

Token buckets cannot be normalized reliably from every provider or client.
When details are missing, report the raw usage evidence and cash/budget effects;
do not invent the missing split. Keep cache-read, cache-write, input and output
rates distinct if reconstructing a cost, and reconcile the estimate with the
provider's billing record.

## Results

| Configuration | Completed / attempted | Total metered usage amount | Total external tool cost | Elapsed time: median / range | Human interventions | Evidence |
| --- | --- | --- | --- | --- | --- | --- |
| [configuration] | [count / count] | [currency + amount or unknown] | [currency + amount or unknown] | [seconds] | [count] | [sanitized reference] |

- Total spend per accepted task: total attributable spend for **all** attempted
  runs divided by the number accepted. If none were accepted, report undefined.
  Explain whether spend means metered charges or allocated subscription cash.
- Publish failures and their costs alongside successes. Show range and sample
  size; avoid precise percentile claims from tiny samples.
- Quality observations beyond tests: [correctness, maintainability, regressions]
- Main failure modes: [tool errors, context limits, budget stop, timeout, other]
- Any manual exclusions, with a reason and their original costs:

## Subscription and monthly fit

Complete this separately from per-run usage. Included credit is not cash paid.

| Item | Value |
| --- | --- |
| Monthly subscription price, currency, tax treatment | |
| Included monthly budget and its unit | |
| Usage windows / reset dates / expiration / carryover | |
| Hard stop, optional overages, or automatic overages | |
| Optional pacing and its observed effect | |
| Total cash paid in the measured period, including upgrades | |
| Free credits / discounts used | |
| Budget consumed in the measured period | |
| Expected task mix and task count next month | |
| Projected budget needed, with uncertainty | |
| Evidence that the sample represents normal usage | |

For a fully observed subscription period, cash per accepted task is that
period's attributable subscription/upgrade cash plus separately billed tool
costs, divided by accepted tasks. Explain shared-plan allocation. For a partial
period, label any monthly allocation or projection as an estimate. A larger
included budget does not establish savings unless the workload fits and the
outputs are acceptable. Light or irregular usage may favor metered billing.

## Conclusion

- For this task mix, the measured tradeoff was:
- The result might change if:
- What remains unknown:
- Next measurement that would change the decision:

## Method references

Check the specific provider and agent versions you used. These primary sources
explain why the distinctions above matter; this template does not reproduce
their mutable price tables.

- [OpenCode provider configuration](https://opencode.ai/docs/providers/)
- [Cline OpenAI-compatible provider settings](https://docs.cline.bot/provider-config/openai-compatible)
- [Anthropic prompt caching documentation](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)

This report contains your measurements. It is not independently verified merely
because it uses this template.
