agent-swarm.devagent-swarm.dev
Guides

Cost & context computation

How cost and context-window numbers are computed across harness providers, and how to read the costSource / contextFormula badges in the UI.

The swarm tracks two related but separate numbers for every model run:

  1. Cost (USD). Each adapter writes one session_costs row per CLI invocation. The API may recompute it from the seeded pricing table.
  2. Context-window usage. Each adapter emits context_usage events; the API persists snapshots and updates aggregate columns on agent_tasks.

This page is the single source of truth for how both numbers are produced.

New session_costs rows keep two USD values:

  • harnessCostUsd is the adapter-reported number. It is useful for comparison, but advisory: adapters can have stale local rates or incomplete usage.
  • totalCostUsd is the API's canonical stored total. For priced rows, the API recomputes it from the active server-side pricing table at the row timestamp; for the other costSource paths below, it intentionally equals the harness report.

How cost is computed

Cost flows through three layers, each annotated with the path's costSource enum value (the dashboard renders this as a small badge next to every cost).

1. Adapter (worker-local)

Every adapter emits a CostData event with its advisory local total, token breakdowns (inputTokens, cacheReadTokens, cacheWriteTokens, outputTokens, reasoningOutputTokens, thinkingTokens), and a provider tag. The dollar value comes from whatever the harness reports — Claude's stream-json carries it directly; Codex doesn't, so the adapter computes locally via computeCodexCostUsd (src/providers/codex-models.ts); pi-ai self-reports stats.cost; etc.

The adapter writes via POST /api/session-costs with the provider field set.

2. API recompute

When the API receives a POST /api/session-costs with a provider tag, it does a synchronous lookup against the pricing table for (provider, model, token_class) at the row's createdAt. Four outcomes:

OutcomecostSourceStored totalCostUsd
A tagged provider/model has the required pricing rows'pricing-table'Canonical server recompute
No provider tag supplied (legacy caller)'harness'Adapter report
A tag was supplied but pricing cannot be completed'unpriced'Adapter report
amp with a billed cost from amp threads usage'harness'What Amp billed for the thread; the token price fills the per-model breakdown
amp without a per-model breakdown (the thread export failed)'estimated'Stream token totals priced at an assumed model
acp whose target reported a USD cost in usage_update'harness'The target's cumulative session cost

harnessCostUsd always preserves the adapter's submitted totalCostUsd, even when the server replaces totalCostUsd with the recomputed value.

Token classes, cache TTLs, and input semantics

Cached reads and cache creation use their own pricing token classes. For Anthropic-billed models, cache_write is the 5-minute creation class (1.25× the base input rate) and cache_write_1h is the one-hour class (2× the base input rate). Claude reports both TTL totals. When a modelBreakdown entry does not have its own TTL split, the recompute distributes its writes using the session's 5m/1h ratio; this is an approximation for sidechains, not a claim that each sidechain independently exposed TTL data.

Input counts have provider-specific cache-read semantics:

Provider familyMeaning of reported inputTokensUncached-input calculation
claude, claude-managed, pi, opencode, dsh, amp, grokExcludes cache readsinputTokens as reported
codexIncludes cache readsmax(0, inputTokens - cacheReadTokens)
cursorExcludes cache reads and writes (the adapter subtracts both, clamped to zero; Cursor's own count includes both)inputTokens as reported

Although OpenCode can route to OpenAI models, the currently shipped recompute treats its input as disjoint from cache reads; production event evidence showed that subtracting cache reads would zero most OpenCode input.

Per-model and request-priced usage

Claude's final result.modelUsage is preserved as modelBreakdown, including sidechain/subagent entries. The API prices every entry at that entry's own model rate, sums those totals, and stores the per-model computed costUsd in the breakdown. Top-level stored token totals are the corresponding breakdown sums when one exists, rather than only the main-thread usage. Because the breakdown takes precedence, the adapter refuses to zero-fill it: a missing, non-finite, or negative token counter on any entry drops the whole breakdown and the session falls back to top-level usage (advisory fields like webSearchRequests and per-model costUSD degrade per-field instead).

web_search is a request-priced class, not a token rate. The manual rows for claude and claude-managed encode Anthropic's $10 per 1,000 requests ($0.01/request); request counts in a model breakdown are added to that model's token cost. A missing web-search rate is treated as $0 so a small search fee cannot discard an otherwise complete token recompute.

claude-managed also adds its $0.08/session-hour runtime_hour fee during the server recompute, using the session duration and the manual runtime_hour pricing row.

3. UI badge

The task-detail and task-detail-sheet views render the costSource next to every cost via <CostSourceBadge> (PRICED, HARNESS, NO RATE, ESTIMATED). Mixed sources within a task aggregate render as HARNESS (the weakest claim).

When both USD values are available and differ, the badge tooltip shows harness and recomputed numbers. The task cost view adds a visible Δ hint above 2%. The OpenTelemetry counter agentswarm.cost.drift.usd records the absolute non-zero difference with a drift_sign attribute, making it the operational watchdog for stale adapter-local pricing or recompute drift.

Attribution coverage and autonomous work

The usage summary separates total spend from spend that could truthfully be assigned to a person. attributableCostUsd is totalCostUsd minus the cost of structurally human-free work. Attribution coverage is therefore attributedCostUsd / attributableCostUsd, not attributed cost divided by all spend. The response also exposes excludedCostUsd and excludedTaskCount so the autonomous population remains visible rather than disappearing from the report.

A task is structurally human-free when its stored task type is heartbeat, heartbeat-checklist, or boot-triage; when its stored JSON tags contain the heartbeat tag (including legacy rows found by the tags LIKE check); or when it is launched by a schedule with no human creator. The schedule rule covers both direct scheduled tasks and workflow roots whose run records that creatorless schedule in its trigger data. Requester-less system follow-ups of requester-less parents are also autonomous.

That classification follows the task tree recursively while descendants have no human requester. This keeps autonomous fan-out out of the denominator, but an explicitly attributed child is treated as a human handoff and stops the classification along that branch. Structurally human-free rows are excluded from attributedCostUsd even if an old or inherited requester id remains on the row, keeping the numerator and denominator a consistent partition.

The classification is computed once, when the task is created, and stored in agent_tasks.isHumanFree (migration 182 backfilled existing tasks). The usage reports read that column instead of walking every task tree per request. Every input to the rule is fixed at creation except where a mutation rewrites one: deleting a user without a replacement clears the requester and the workflow run creator, deleting a workflow removes its runs, deleting a task removes a parent, and completing a task can add tags. Each of those reclassifies the affected tasks and their descendants, so the stored flag matches the rule over the current rows.

The work per request is bounded, because the size of a task tree is not. A request reclassifies at most 500 tasks, walking down from the changed task breadth first, in the same transaction as the change. Whatever is left of a larger tree is recorded in human_free_reclassify_queue in that transaction, and a background drain finishes it in batches of 500, one transaction each. Until the queue is empty the stored flag of a queued descendant is its previous value, so the usage reports can lag by seconds for a very large tree. The queue survives a restart.

The usage endpoints (/api/session-costs/summary, /api/attribution/by-person) cache each reply per filter set. A reply is fresh for 30 seconds. Until its total age reaches 120 seconds (counted from when it was loaded), requests still get it while one background query refreshes it. A request that finds it older than that waits for a new query. The first request past the 30 seconds only starts the refresh and still receives the old reply, so a new session shows up on the request after that: about one poll later for a client that polls every 30 seconds. Changing a credential's plan or name clears the cache.

The By Person view uses the same requester data model, but it is not a grouping of the cost denominator and reads the same stored flag. It reports work outcomes rather than a cost score: human-requested root tasks supply Problems Initiated and Problems Shipped, while each person's full task trees supply Agents Reached, Repos Reached, and Surfaces Reached. Requester-less autonomous roots and heartbeat-classified roots do not belong to a person. The metrics stay side by side and are never summed into a composite ranking.

How the pricing table is populated

The pricing table starts with a boot seed from the vendored models.dev snapshot at src/be/modelsdev-cache.json plus a small set of manual overrides for items models.dev doesn't carry. The committed snapshot is now fallback-only for pricing freshness: it gives a cold-start DB usable rows when models.dev is unavailable, while src/be/pricing-refresh.ts owns live price updates.

After boot, the API server runs an in-process models.dev refresher once immediately and then every 12 hours. It fetches https://models.dev/api.json with If-None-Match, projects the response through the same buildModelsDevSeedRows() logic, inserts a new effective row only for new models or changed prices, and prunes history to the latest two rows per (provider, model, token_class).

The same refresh also feeds the runtime model catalog at GET /api/models-catalog (src/be/models-catalog.ts) — a slim projection of the picker-reachable providers. The UI model picker prefers that live catalog, so newly released models appear without redeploying; the committed snapshot (still symlinked at ui/src/lib/modelsdev-cache.json) remains the build-time fallback for names, labels, and context windows while the request is in flight or the server predates the endpoint.

  • Projection rules live in src/be/seed-pricing.ts:
    • Anthropic models → rows under both provider='claude' AND provider='claude-managed'. Shortnames (opus/sonnet/haiku/fable/mythos) also land under the current default full id; Fable 5.1 and Mythos 5.1 use verified fallback rates until the vendored models.dev snapshot includes them.
    • OpenAI models → provider='codex'.
    • OpenRouter models → provider='opencode', provider='pi', provider='dsh' and provider='grok'; google/* models also land under provider='gemini'.
    • DeepSeek direct-API models (the models.dev deepseek section, bare ids such as deepseek-v4-pro) → provider='dsh'.
    • Anthropic, OpenAI, Google and Fireworks models (the models.dev anthropic, openai, google and fireworks-ai sections, in the vendor's own ids) → provider='amp'. Amp reports gpt-5-nano-2025-08-07 and accounts/fireworks/models/glm-5p3-flash; the lookup drops -YYYY-MM-DD snapshot dates and a provider/ pin prefix.
    • xAI models (the models.dev xai section, bare ids such as grok-4.6) → provider='grok'. The base rates only: xAI's higher rate above 200k context is not applied. They price the per-model breakdown; the row's total is what xAI billed (see grok below).
    • OpenAI, Anthropic, Google and xAI models (bare vendor ids) → provider='cursor'. Cursor bills the vendor's API rates and names the models by the vendor's id.
    • Cursor's own Composer models → provider='cursor' from CURSOR_FIRST_PARTY_PRICING, at Cursor's published Composer 2.5 (Fast) rates: $3.00 input, $0.50 cache read, $15.00 output per million tokens. Fast is Composer's default variant, and composer-2 is retired and rerouted to Composer 2.5. default (Auto) bills at the routed model's price, has no rate of its own, and stays unpriced. See Cursor.
  • Manual overrides (claude-managed runtime_hour at $0.08/hr, devin acu at $2.25): MANUAL_PRICING_OVERRIDES in the same file. Each entry carries its source URL and a verified date.
  • Runtime refresh: src/be/pricing-refresh.ts updates pricing rows in-place after boot and every 12 hours. It only adds newer effective rows and never deletes pinned entries from the committed snapshot.
  • Snapshot refresh procedure: run bun run scripts/refresh-modelsdev-pricing.ts when the committed fallback/UI catalog needs a source update. Commit it alongside the PR.

Operator reference: src/providers/pricing-sources.md.

How context-window usage is computed

The unified formula

After Phase 9, every adapter uses one formula:

contextUsedTokens = inputTokens + cacheReadTokens + cacheCreateTokens + outputTokens

Helpers: computeContextUsedUnified and clampContextPercent in src/utils/context-window.ts. The emitted event carries contextFormula: 'input-cache-output'.

Pi-mono is the exception: pi-ai owns the formula and we just relay its numbers. Those snapshots are tagged contextFormula: 'pi-delegated'. Devin's API doesn't report context info at all; we omit the event rather than fake zeros.

Per-model window resolution

getContextWindowSize(model) resolves:

  • Shortnames (opus/sonnet/haiku/fable/mythos)
  • Family-versioned ids (claude-sonnet-4-6)
  • Claude 5.1 premium ids (claude-fable-5-1/claude-mythos-5-1)
  • Dated full ids (claude-sonnet-4-6-20251004) — by stripping the 8-digit date suffix and retrying

Fallback is 200k. Pre-Phase 4 the dated form fell to 200k unconditionally — wildly wrong for opus/sonnet 4.x.

peakContextTokens and contextWindowSize

agent_tasks.peakContextTokens (renamed from totalContextTokensUsed in migration 063) is a monotonic max across all snapshots for the task — never regresses when a later snapshot reports a smaller value. This mirrors Claude Code's status-line "peak context" idea.

agent_tasks.contextWindowSize is set on the FIRST snapshot that carries one, not gated on eventType='completion'. Subsequent snapshots leave it alone.

Per-provider notes

  • claude / claude-managed: token rates from models.dev. claude-managed also has a per-session-hour runtime fee (token_class='runtime_hour'); the worker computes a preview locally via claude-managed-pricing.ts, and the API's recompute path overrides with the canonical value.
  • codex: each thread/tokenUsage/updated notification carries total (cumulative for the thread), last (the most recent model request), and modelContextWindow. On each terminal turn, including failed and interrupted turns, the adapter converts the latest total to a delta before it updates CostData. This prevents the same usage from being charged again after a later turn. Context comes from last on every notification, so it updates during the turn: last.inputTokens (which already includes cached input) plus last.outputTokens, against modelContextWindow when Codex reports one and the models.dev window otherwise. A turn's summed usage is never reported as context, because one turn can span many model requests and tool calls. Cache writes are input details. Codex 0.153.4 reports cache-write tokens separately, and the adapter includes them in CostData.
  • pi-mono: cost passes through verbatim from pi-ai's stats.cost. Context snapshots tag contextFormula: 'pi-delegated'. durationMs is now real wallclock (was hardcoded 0). Per-turn outputTokens are derived from session-stats delta.
  • opencode: passthrough through OpenRouter. The unified formula applies; contextPercent is clamped to [0, 100].
  • dsh: dsh reports tokens, not money. Each status.step_end.usage (one model call) becomes a context snapshot with the unified formula; the summed tokens become CostData with totalCostUsd: 0 and provider: 'dsh', and the API prices them from the dsh rows (openrouter/ stripped for OpenRouter models). dsh's own deepseek-flash id is not in models.dev, so it stays unpriced.
  • amp: Amp's stream reports tokens but no money and no model. Each assistant message's usage becomes a context snapshot with the unified formula (window unknown until the session ends). At the end the adapter reads amp threads usage <id> --details for what Amp billed and amp threads export <id> for the model of every request and Amp's context window. When every request was billed through Amp, the billed dollars go out as totalCostUsd and the API keeps them as costSource: 'harness': they cover subagent threads and title requests the export misses (on short live threads the token price was 30 to 70% of the billed figure). Otherwise totalCostUsd is 0 and the API prices the per-model breakdown; a request routed through a linked provider such as a ChatGPT subscription records $0 in Amp and is billed outside it, so the token price is the conservative figure. Amp's cache_creation_input_tokens is billed as cache writes for Anthropic models and as plain input for every other vendor. When the export fails, the adapter posts the stream's raw counts under the requested mode or pin with no breakdown. The API never records $0 for it, because budget admission sums totalCostUsd and a zero row would let repeated sessions spend unchecked. It prices the stream totals at an assumed model and tags the row estimated: the pin when the table has its rates, otherwise the model the mode runs (AMP_ESTIMATE_MODELS in src/utils/amp-models.ts, measured 2026-10-05: low → GLM-5.3 Flash, medium → Claude Opus 5.5, high → GPT-6 Astra, ultra → Claude Fable 5.1), otherwise the medium model. Cache-creation tokens follow the same Anthropic-only cache-write rule. Only a broken pricing seed (no rates for the fallback model) leaves the row unpriced.
  • grok: the Grok CLI runs on the shared ACP client but sends no ACP usage. The prompt response's _meta.usage carries the prompt's cumulative tokens and costUsdTicks (xAI's bill, 1e-10 USD). Grok's inputTokens include cache reads and its outputTokens exclude reasoning (verified live: the converted tokens at models.dev rates equal the ticks exactly), so the adapter reports input − cacheRead and output + reasoning. The billed USD wins over the table (costSource: 'harness'), as for amp; the grok rows fill the breakdown, one entry per _meta.usage.modelUsage model. The row's model and its context window are _meta.modelId, the model that ran, not the one requested. An openrouter/<id> model reports no ticks and prices from the OpenRouter rows, each reported model keeping the openrouter/ prefix. An abort waits up to 3s for the cancelled prompt's answer, so a session the server already closed keeps its cost row.
  • cursor: @cursor/sdk reports tokens per run, summed over the run's model calls, and no per-call figure. agent.getUsage() (billed cents) answers feature_unavailable for local agents, so CostData carries tokens with totalCostUsd: 0 and provider: 'cursor', and the API prices them from the cursor rows. Cursor's inputTokens includes cache reads and writes; the adapter subtracts both, clamped to zero, to report fresh input separately. The context snapshot is the run's per-call average ((input + cacheWrite + output) / (tool rounds + 1)), tagged peak-proxy because it under-reads the last call. Composer models price from Cursor's published rates (see above); default (Auto) stays unpriced. The result is a list-price estimate: it ignores plan discounts, included usage, and the Teams/Enterprise Cursor Token Rate.
  • devin: ACU-based pricing (token_class='acu', $2.25 per ACU). No per-token cost. No context events (the API doesn't report context info) — peakContextTokens remains null for devin tasks.

Gotchas & known limitations

  • Internal-ai Gemini calls are not yet costed. src/utils/internal-ai/models.ts:19-25 routes through OpenRouter for summarization/rating but doesn't yet emit session_costs rows. The pricing table now has gemini rows ready; instrumentation is a follow-up.
  • Codex context is the latest request, not the turn. Cost uses per-turn deltas of the cumulative total; context uses the per-request last snapshot. Rows written before this split summed a whole turn and could exceed the window (clamped to 100%). Old peak-proxy rows in task_context_snapshots remain correct for their original formula and are not directly comparable to current input-cache-output rows.
  • Model-id key mismatch. Some adapters use harness-prefixed ids (openai-codex/gpt-5.4-mini); pricing-table seeds use the stripped form (gpt-5.4-mini). Pick one convention if you're adding a new mapping.
  • Timestamp convention split. session_costs.createdAt and task_context_snapshots.createdAt are TEXT ISO 8601; pricing.effective_from / budgets.createdAt are INTEGER epoch-ms. Documented in 046_budgets_and_pricing.sql:17-22; not a near-term cleanup.

ACP prompt usage

ACP targets can supply optional token usage in the session/prompt response. The adapter maps inputTokens and outputTokens directly, and maps cachedReadTokens / cachedWriteTokens to the corresponding cache counters. Missing counts remain undefined at the adapter boundary. JSON serialization omits them; the session-cost API and DB deliberately retain their existing zero-coalescing behavior for input, output, cache-read, and cache-write counters. Persisted costs, pricing computation, and UI displays therefore cannot distinguish absent usage from measured zero. Preserving that distinction end to end would require a separate request/DB/response contract change. Context-window usage_update totals never substitute for billing tokens.

The acp_prompt_response raw session log retains the bounded, scrubbed usage payload and optional _meta; usage: null records an absent payload.

A target can report money in usage_update.cost (amount plus an ISO 4217 currency). ACP defines amount as the session's cumulative cost, and OpenCode 1.18.34 sends it that way: two turns in one session reported $0.0172 and then $0.0345. The swarm sends one prompt per session, so the adapter takes the latest USD amount as the session total and sends it as totalCostUsd. The API stores it as costSource: 'harness'.

ACP does not infer USD or a pricing alias from its target. When the target reports no cost, or reports it in a currency other than USD, totalCostUsd stays 0 and a session with an unresolved (acp, model) pricing identity remains unpriced, even when token counts are available. The raw usage_update session log keeps the reported value. A requested model ID is insufficient evidence for a rate mapping when the target can reject that model and fall back to another; a target-reported cost does not depend on the model ID, so it needs no rate.

On this page