Cost & context computation
How cost and context-window numbers are computed across harness providers, and how to read the costSource / contextFormula badges in the UI.
The swarm tracks two related but separate numbers for every model run:
- Cost (USD). Each adapter writes one
session_costsrow per CLI invocation. The API may recompute it from the seeded pricing table. - Context-window usage. Each adapter emits
context_usageevents; the API persists snapshots and updates aggregate columns onagent_tasks.
This page is the single source of truth for how both numbers are produced.
New session_costs rows keep two USD values:
harnessCostUsdis the adapter-reported number. It is useful for comparison, but advisory: adapters can have stale local rates or incomplete usage.totalCostUsdis the API's canonical stored total. For priced rows, the API recomputes it from the active server-side pricing table at the row timestamp; for the othercostSourcepaths below, it intentionally equals the harness report.
How cost is computed
Cost flows through three layers, each annotated with the path's costSource enum value (the dashboard renders this as a small badge next to every cost).
1. Adapter (worker-local)
Every adapter emits a CostData event with its advisory local total, token breakdowns (inputTokens, cacheReadTokens, cacheWriteTokens, outputTokens, reasoningOutputTokens, thinkingTokens), and a provider tag. The dollar value comes from whatever the harness reports — Claude's stream-json carries it directly; Codex doesn't, so the adapter computes locally via computeCodexCostUsd (src/providers/codex-models.ts); pi-ai self-reports stats.cost; etc.
The adapter writes via POST /api/session-costs with the provider field set.
2. API recompute
When the API receives a POST /api/session-costs with a provider tag, it does a synchronous lookup against the pricing table for (provider, model, token_class) at the row's createdAt. Four outcomes:
| Outcome | costSource | Stored totalCostUsd |
|---|---|---|
| A tagged provider/model has the required pricing rows | 'pricing-table' | Canonical server recompute |
No provider tag supplied (legacy caller) | 'harness' | Adapter report |
| A tag was supplied but pricing cannot be completed | 'unpriced' | Adapter report |
amp with a billed cost from amp threads usage | 'harness' | What Amp billed for the thread; the token price fills the per-model breakdown |
amp without a per-model breakdown (the thread export failed) | 'estimated' | Stream token totals priced at an assumed model |
acp whose target reported a USD cost in usage_update | 'harness' | The target's cumulative session cost |
harnessCostUsd always preserves the adapter's submitted totalCostUsd, even
when the server replaces totalCostUsd with the recomputed value.
Token classes, cache TTLs, and input semantics
Cached reads and cache creation use their own pricing token classes. For
Anthropic-billed models, cache_write is the 5-minute creation class (1.25×
the base input rate) and cache_write_1h is the one-hour class (2× the base
input rate). Claude reports both TTL totals. When a modelBreakdown entry
does not have its own TTL split, the recompute distributes its writes using the
session's 5m/1h ratio; this is an approximation for sidechains, not a claim
that each sidechain independently exposed TTL data.
Input counts have provider-specific cache-read semantics:
| Provider family | Meaning of reported inputTokens | Uncached-input calculation |
|---|---|---|
claude, claude-managed, pi, opencode, dsh, amp, grok | Excludes cache reads | inputTokens as reported |
codex | Includes cache reads | max(0, inputTokens - cacheReadTokens) |
cursor | Excludes cache reads and writes (the adapter subtracts both, clamped to zero; Cursor's own count includes both) | inputTokens as reported |
Although OpenCode can route to OpenAI models, the currently shipped recompute treats its input as disjoint from cache reads; production event evidence showed that subtracting cache reads would zero most OpenCode input.
Per-model and request-priced usage
Claude's final result.modelUsage is preserved as modelBreakdown, including
sidechain/subagent entries. The API prices every entry at that entry's own
model rate, sums those totals, and stores the per-model computed costUsd in
the breakdown. Top-level stored token totals are the corresponding breakdown
sums when one exists, rather than only the main-thread usage. Because the
breakdown takes precedence, the adapter refuses to zero-fill it: a missing,
non-finite, or negative token counter on any entry drops the whole breakdown
and the session falls back to top-level usage (advisory fields like
webSearchRequests and per-model costUSD degrade per-field instead).
web_search is a request-priced class, not a token rate. The manual rows for
claude and claude-managed encode Anthropic's $10 per 1,000 requests
($0.01/request); request counts in a model breakdown are added to that
model's token cost. A missing web-search rate is treated as $0 so a small
search fee cannot discard an otherwise complete token recompute.
claude-managed also adds its $0.08/session-hour runtime_hour fee during
the server recompute, using the session duration and the manual runtime_hour
pricing row.
3. UI badge
The task-detail and task-detail-sheet views render the costSource next to every cost via <CostSourceBadge> (PRICED, HARNESS, NO RATE, ESTIMATED). Mixed sources within a task aggregate render as HARNESS (the weakest claim).
When both USD values are available and differ, the badge tooltip shows harness
and recomputed numbers. The task cost view adds a visible Δ hint above 2%.
The OpenTelemetry counter agentswarm.cost.drift.usd records the absolute
non-zero difference with a drift_sign attribute, making it the operational
watchdog for stale adapter-local pricing or recompute drift.
Attribution coverage and autonomous work
The usage summary separates total spend from spend that could truthfully be
assigned to a person. attributableCostUsd is totalCostUsd minus the cost of
structurally human-free work. Attribution coverage is therefore
attributedCostUsd / attributableCostUsd, not attributed cost divided by all
spend. The response also exposes excludedCostUsd and excludedTaskCount so
the autonomous population remains visible rather than disappearing from the
report.
A task is structurally human-free when its stored task type is heartbeat,
heartbeat-checklist, or boot-triage; when its stored JSON tags contain the
heartbeat tag (including legacy rows found by the tags LIKE check); or when
it is launched by a schedule with no human creator. The schedule rule covers
both direct scheduled tasks and workflow roots whose run records that
creatorless schedule in its trigger data. Requester-less system follow-ups of
requester-less parents are also autonomous.
That classification follows the task tree recursively while descendants have
no human requester. This keeps autonomous fan-out out of the denominator, but
an explicitly attributed child is treated as a human handoff and stops the
classification along that branch. Structurally human-free rows are excluded
from attributedCostUsd even if an old or inherited requester id remains on
the row, keeping the numerator and denominator a consistent partition.
The classification is computed once, when the task is created, and stored in
agent_tasks.isHumanFree (migration 182 backfilled existing tasks). The usage
reports read that column instead of walking every task tree per request. Every
input to the rule is fixed at creation except where a mutation rewrites one:
deleting a user without a replacement clears the requester and the workflow run
creator, deleting a workflow removes its runs, deleting a task removes a parent,
and completing a task can add tags. Each of those reclassifies the affected tasks
and their descendants, so the stored flag matches the rule over the current rows.
The work per request is bounded, because the size of a task tree is not. A
request reclassifies at most 500 tasks, walking down from the changed task
breadth first, in the same transaction as the change. Whatever is left of a
larger tree is recorded in human_free_reclassify_queue in that transaction, and
a background drain finishes it in batches of 500, one transaction each. Until the
queue is empty the stored flag of a queued descendant is its previous value, so
the usage reports can lag by seconds for a very large tree. The queue survives a
restart.
The usage endpoints (/api/session-costs/summary, /api/attribution/by-person)
cache each reply per filter set. A reply is fresh for 30 seconds. Until its total
age reaches 120 seconds (counted from when it was loaded), requests still get it
while one background query refreshes it. A request that finds it older than that
waits for a new query. The first request past the 30 seconds only starts the
refresh and still receives the old reply, so a new session shows up on the
request after that: about one poll later for a client that polls every 30
seconds. Changing a credential's plan or name clears the cache.
The By Person view uses the same requester data model, but it is not a grouping of the cost denominator and reads the same stored flag. It reports work outcomes rather than a cost score: human-requested root tasks supply Problems Initiated and Problems Shipped, while each person's full task trees supply Agents Reached, Repos Reached, and Surfaces Reached. Requester-less autonomous roots and heartbeat-classified roots do not belong to a person. The metrics stay side by side and are never summed into a composite ranking.
How the pricing table is populated
The pricing table starts with a boot seed from the vendored models.dev snapshot at src/be/modelsdev-cache.json plus a small set of manual overrides for items models.dev doesn't carry. The committed snapshot is now fallback-only for pricing freshness: it gives a cold-start DB usable rows when models.dev is unavailable, while src/be/pricing-refresh.ts owns live price updates.
After boot, the API server runs an in-process models.dev refresher once immediately and then every 12 hours. It fetches https://models.dev/api.json with If-None-Match, projects the response through the same buildModelsDevSeedRows() logic, inserts a new effective row only for new models or changed prices, and prunes history to the latest two rows per (provider, model, token_class).
The same refresh also feeds the runtime model catalog at GET /api/models-catalog (src/be/models-catalog.ts) — a slim projection of the picker-reachable providers. The UI model picker prefers that live catalog, so newly released models appear without redeploying; the committed snapshot (still symlinked at ui/src/lib/modelsdev-cache.json) remains the build-time fallback for names, labels, and context windows while the request is in flight or the server predates the endpoint.
- Projection rules live in
src/be/seed-pricing.ts:- Anthropic models → rows under both
provider='claude'ANDprovider='claude-managed'. Shortnames (opus/sonnet/haiku/fable/mythos) also land under the current default full id; Fable 5.1 and Mythos 5.1 use verified fallback rates until the vendored models.dev snapshot includes them. - OpenAI models →
provider='codex'. - OpenRouter models →
provider='opencode',provider='pi',provider='dsh'andprovider='grok';google/*models also land underprovider='gemini'. - DeepSeek direct-API models (the models.dev
deepseeksection, bare ids such asdeepseek-v4-pro) →provider='dsh'. - Anthropic, OpenAI, Google and Fireworks models (the models.dev
anthropic,openai,googleandfireworks-aisections, in the vendor's own ids) →provider='amp'. Amp reportsgpt-5-nano-2025-08-07andaccounts/fireworks/models/glm-5p3-flash; the lookup drops-YYYY-MM-DDsnapshot dates and aprovider/pin prefix. - xAI models (the models.dev
xaisection, bare ids such asgrok-4.6) →provider='grok'. The base rates only: xAI's higher rate above 200k context is not applied. They price the per-model breakdown; the row's total is what xAI billed (see grok below). - OpenAI, Anthropic, Google and xAI models (bare vendor ids) →
provider='cursor'. Cursor bills the vendor's API rates and names the models by the vendor's id. - Cursor's own Composer models →
provider='cursor'fromCURSOR_FIRST_PARTY_PRICING, at Cursor's published Composer 2.5 (Fast) rates: $3.00 input, $0.50 cache read, $15.00 output per million tokens. Fast is Composer's default variant, andcomposer-2is retired and rerouted to Composer 2.5.default(Auto) bills at the routed model's price, has no rate of its own, and staysunpriced. See Cursor.
- Anthropic models → rows under both
- Manual overrides (claude-managed
runtime_hourat $0.08/hr, devinacuat $2.25):MANUAL_PRICING_OVERRIDESin the same file. Each entry carries its source URL and averifieddate. - Runtime refresh:
src/be/pricing-refresh.tsupdates pricing rows in-place after boot and every 12 hours. It only adds newer effective rows and never deletes pinned entries from the committed snapshot. - Snapshot refresh procedure: run
bun run scripts/refresh-modelsdev-pricing.tswhen the committed fallback/UI catalog needs a source update. Commit it alongside the PR.
Operator reference: src/providers/pricing-sources.md.
How context-window usage is computed
The unified formula
After Phase 9, every adapter uses one formula:
contextUsedTokens = inputTokens + cacheReadTokens + cacheCreateTokens + outputTokensHelpers: computeContextUsedUnified and clampContextPercent in src/utils/context-window.ts. The emitted event carries contextFormula: 'input-cache-output'.
Pi-mono is the exception: pi-ai owns the formula and we just relay its numbers. Those snapshots are tagged contextFormula: 'pi-delegated'. Devin's API doesn't report context info at all; we omit the event rather than fake zeros.
Per-model window resolution
getContextWindowSize(model) resolves:
- Shortnames (
opus/sonnet/haiku/fable/mythos) - Family-versioned ids (
claude-sonnet-4-6) - Claude 5.1 premium ids (
claude-fable-5-1/claude-mythos-5-1) - Dated full ids (
claude-sonnet-4-6-20251004) — by stripping the 8-digit date suffix and retrying
Fallback is 200k. Pre-Phase 4 the dated form fell to 200k unconditionally — wildly wrong for opus/sonnet 4.x.
peakContextTokens and contextWindowSize
agent_tasks.peakContextTokens (renamed from totalContextTokensUsed in migration 063) is a monotonic max across all snapshots for the task — never regresses when a later snapshot reports a smaller value. This mirrors Claude Code's status-line "peak context" idea.
agent_tasks.contextWindowSize is set on the FIRST snapshot that carries one, not gated on eventType='completion'. Subsequent snapshots leave it alone.
Per-provider notes
- claude / claude-managed: token rates from models.dev. claude-managed also has a per-session-hour runtime fee (
token_class='runtime_hour'); the worker computes a preview locally viaclaude-managed-pricing.ts, and the API's recompute path overrides with the canonical value. - codex: each
thread/tokenUsage/updatednotification carriestotal(cumulative for the thread),last(the most recent model request), andmodelContextWindow. On each terminal turn, including failed and interrupted turns, the adapter converts the latesttotalto a delta before it updatesCostData. This prevents the same usage from being charged again after a later turn. Context comes fromlaston every notification, so it updates during the turn:last.inputTokens(which already includes cached input) pluslast.outputTokens, againstmodelContextWindowwhen Codex reports one and the models.dev window otherwise. A turn's summed usage is never reported as context, because one turn can span many model requests and tool calls. Cache writes are input details. Codex 0.153.4 reports cache-write tokens separately, and the adapter includes them inCostData. - pi-mono: cost passes through verbatim from pi-ai's
stats.cost. Context snapshots tagcontextFormula: 'pi-delegated'.durationMsis now real wallclock (was hardcoded 0). Per-turnoutputTokensare derived from session-stats delta. - opencode: passthrough through OpenRouter. The unified formula applies;
contextPercentis clamped to [0, 100]. - dsh: dsh reports tokens, not money. Each
status.step_end.usage(one model call) becomes a context snapshot with the unified formula; the summed tokens becomeCostDatawithtotalCostUsd: 0andprovider: 'dsh', and the API prices them from thedshrows (openrouter/stripped for OpenRouter models). dsh's owndeepseek-flashid is not in models.dev, so it staysunpriced. - amp: Amp's stream reports tokens but no money and no model. Each assistant message's usage becomes a context snapshot with the unified formula (window unknown until the session ends). At the end the adapter reads
amp threads usage <id> --detailsfor what Amp billed andamp threads export <id>for the model of every request and Amp's context window. When every request was billed through Amp, the billed dollars go out astotalCostUsdand the API keeps them ascostSource: 'harness': they cover subagent threads and title requests the export misses (on short live threads the token price was 30 to 70% of the billed figure). OtherwisetotalCostUsdis 0 and the API prices the per-model breakdown; a request routed through a linked provider such as a ChatGPT subscription records $0 in Amp and is billed outside it, so the token price is the conservative figure. Amp'scache_creation_input_tokensis billed as cache writes for Anthropic models and as plain input for every other vendor. When the export fails, the adapter posts the stream's raw counts under the requested mode or pin with no breakdown. The API never records $0 for it, because budget admission sumstotalCostUsdand a zero row would let repeated sessions spend unchecked. It prices the stream totals at an assumed model and tags the rowestimated: the pin when the table has its rates, otherwise the model the mode runs (AMP_ESTIMATE_MODELSinsrc/utils/amp-models.ts, measured 2026-10-05:low→ GLM-5.3 Flash,medium→ Claude Opus 5.5,high→ GPT-6 Astra,ultra→ Claude Fable 5.1), otherwise themediummodel. Cache-creation tokens follow the same Anthropic-only cache-write rule. Only a broken pricing seed (no rates for the fallback model) leaves the rowunpriced. - grok: the Grok CLI runs on the shared ACP client but sends no ACP
usage. The prompt response's_meta.usagecarries the prompt's cumulative tokens andcostUsdTicks(xAI's bill, 1e-10 USD). Grok'sinputTokensinclude cache reads and itsoutputTokensexclude reasoning (verified live: the converted tokens at models.dev rates equal the ticks exactly), so the adapter reportsinput − cacheReadandoutput + reasoning. The billed USD wins over the table (costSource: 'harness'), as for amp; thegrokrows fill the breakdown, one entry per_meta.usage.modelUsagemodel. The row's model and its context window are_meta.modelId, the model that ran, not the one requested. Anopenrouter/<id>model reports no ticks and prices from the OpenRouter rows, each reported model keeping theopenrouter/prefix. An abort waits up to 3s for the cancelled prompt's answer, so a session the server already closed keeps its cost row. - cursor:
@cursor/sdkreports tokens per run, summed over the run's model calls, and no per-call figure.agent.getUsage()(billed cents) answersfeature_unavailablefor local agents, soCostDatacarries tokens withtotalCostUsd: 0andprovider: 'cursor', and the API prices them from thecursorrows. Cursor'sinputTokensincludes cache reads and writes; the adapter subtracts both, clamped to zero, to report fresh input separately. The context snapshot is the run's per-call average ((input + cacheWrite + output) / (tool rounds + 1)), taggedpeak-proxybecause it under-reads the last call. Composer models price from Cursor's published rates (see above);default(Auto) staysunpriced. The result is a list-price estimate: it ignores plan discounts, included usage, and the Teams/Enterprise Cursor Token Rate. - devin: ACU-based pricing (
token_class='acu', $2.25 per ACU). No per-token cost. No context events (the API doesn't report context info) —peakContextTokensremainsnullfor devin tasks.
Gotchas & known limitations
- Internal-ai Gemini calls are not yet costed.
src/utils/internal-ai/models.ts:19-25routes through OpenRouter for summarization/rating but doesn't yet emitsession_costsrows. The pricing table now hasgeminirows ready; instrumentation is a follow-up. - Codex context is the latest request, not the turn. Cost uses per-turn deltas of the cumulative
total; context uses the per-requestlastsnapshot. Rows written before this split summed a whole turn and could exceed the window (clamped to 100%). Oldpeak-proxyrows intask_context_snapshotsremain correct for their original formula and are not directly comparable to currentinput-cache-outputrows. - Model-id key mismatch. Some adapters use harness-prefixed ids (
openai-codex/gpt-5.4-mini); pricing-table seeds use the stripped form (gpt-5.4-mini). Pick one convention if you're adding a new mapping. - Timestamp convention split.
session_costs.createdAtandtask_context_snapshots.createdAtare TEXT ISO 8601;pricing.effective_from/budgets.createdAtare INTEGER epoch-ms. Documented in046_budgets_and_pricing.sql:17-22; not a near-term cleanup.
Related docs
- Harness providers — provider-specific quirks
src/providers/pricing-sources.md— operator workflowBUSINESS_USE.md— flow diagrams fortask/agent/apievents
ACP prompt usage
ACP targets can supply optional token usage in the session/prompt response.
The adapter maps inputTokens and outputTokens directly, and maps
cachedReadTokens / cachedWriteTokens to the corresponding cache counters.
Missing counts remain undefined at the adapter boundary. JSON serialization
omits them; the session-cost API and DB deliberately retain their existing
zero-coalescing behavior for input, output, cache-read, and cache-write counters.
Persisted costs, pricing computation, and UI displays therefore cannot distinguish
absent usage from measured zero. Preserving that distinction end to end would
require a separate request/DB/response contract change. Context-window
usage_update totals never substitute for billing tokens.
The acp_prompt_response raw session log retains the bounded, scrubbed usage
payload and optional _meta; usage: null records an absent payload.
A target can report money in usage_update.cost (amount plus an ISO 4217
currency). ACP defines amount as the session's cumulative cost, and OpenCode
1.18.34 sends it that way: two turns in one session reported $0.0172 and then
$0.0345. The swarm sends one prompt per session, so the adapter takes the latest
USD amount as the session total and sends it as totalCostUsd. The API stores it
as costSource: 'harness'.
ACP does not infer USD or a pricing alias from its target. When the target
reports no cost, or reports it in a currency other than USD, totalCostUsd stays
0 and a session with an unresolved (acp, model) pricing identity remains
unpriced, even when token counts are available. The raw usage_update session
log keeps the reported value. A requested model ID is insufficient evidence for
a rate mapping when the target can reject that model and fall back to another; a
target-reported cost does not depend on the model ID, so it needs no rate.