Release Notes — Week of September 28 to October 5, 2026
Five releases let swarms make cheap workflow decisions with human review, run Claude and memory on enterprise clouds, and prove results on a public benchmark, across 190 merged commits.
Five releases shipped this week — v1.158.0 through v1.162.0 — across 190 merged commits. The through-line: swarms got cheaper and safer at deciding, they run where enterprises already run, and we started publishing evidence of how well they work. A note on scope: 14 commits landed on main after the v1.162.0 tag and ship in the next release; each is marked below.
Highlights
System-one decisions with human review
A native system-one-decision workflow node routes branches through a fast, cheap model call — including via OpenRouter — instead of spinning up a full agent task. Routine routing and triage no longer need an agent in the loop. Provider keys are preflighted before the run starts, so a missing key fails at save time, not mid-run. And when the model's confidence lands inside a band you configure, the answer pauses for human review instead of guessing (#1715). Uncertain calls escalate to people by design, which is the answer to "who watches the agents" that actually ships. Setup walkthrough, including the laya provider: (#1750).
Claude on your gateway, Foundry, Bedrock or Vertex
Claude workers now accept gateway, Microsoft Foundry, AWS Bedrock and Google Vertex routes (#1816). Pair them with the Azure / Foundry embeddings preset and Memory integration card so memory stays in the same cloud as the model (#1836). Enterprise swarms no longer route around their existing cloud commitments.
Azure DevOps integration
A minimal Azure DevOps integration: Azure Repos plus PR service hooks (#1812). A follow-up acks partial payloads instead of returning 500 (#1823).
Live model catalog
Model lists come from a live catalog instead of hard-coded tables, tiers resolve when a task is claimed, and a task whose model the assignee harness cannot run is rejected up front instead of failing mid-run (#1691, #1733, #1764).
Delegate from Claude Code with claude-mod
claude-mod puts the swarm pane, a delegate tool, and live task logs inside Claude Code, so you can fan work out to the swarm without leaving your editor (#1838).
Public agent benchmark
The evals app shipped as a public benchmark: a leaderboard with Pareto frontier and ranking table (#1755), scenario cards and a heatmap (#1759), and an analytics API for picking the best setup (#1744). Swarm scenarios — fan-out, worker recovery, implement-review, capability routing, human-in-loop — run against solo baselines (#1745, #1756), with reasoning effort configurable per harness and model (#1710) and refusal-gated publishing (#1761).
Improvements
Workflows
- Standard-webhooks verification — webhook triggers verify
webhook-id/webhook-timestamp/webhook-signatureheaders against thewhsec_signing secret (#1760). - Paged workflow runs —
GET /api/workflows/{id}/runsis paged and list rows drop run context (#1699).
Memory
- Logical paths — memories move and filter under
/longtermvia key, newKey, keyPrefix and path weight (#1814). - Longterm key tree and whole-memory view on
/memory— on main after v1.162.0, ships in the next release (#1856). - Count-based retention for context_versions (#1808).
Slack and extensions
- Native Slack working status for the whole life of an ask (#1770).
- Opt-in inline extension install for leads and operators (#1828).
- Deploy-awareness catalog extension (#1826).
Harnesses and models
- pi — opt-in codemode (#1731) with a
PI_CODEMODE_MODELSflag (#1824), installed MCP servers through pi's MCP extension (#1730), and non-core tools deferred behind tool_search (#1729). - Codex CLI 0.159.0 with GPT-6.1 Sol (#1719).
- OpenCode emits assistant text so outputSchema validation works (#1753).
- Auth resilience — seat-aware Claude OAuth key selection (#1765) and Codex pool login benching after 2 auth failures (#1763).
Dashboard
- Contextual session panel on every page (#1688).
- agent-fs review space in the dashboard, off by default (#1728).
- Reference icon set for task status — on main after v1.162.0, ships in the next release (#1860).
- Dev mode for version gates (#1726).
Telemetry
- schema_version 2 identity envelope (#1780) and dashboard events honor the telemetry opt-out (#1773).
Docs
- Multi-agent orchestration patterns guide (#1864) and team collaboration feature page (#1865) — both on main after v1.162.0, ship in the next release.
- Workflow example configs and routing corrected (#1862).
Bug Fixes
API performance and event-loop stalls
- Slim task lists — SQL projection, opt-in total and a timeline shape for
GET /api/tasks(#1697); opt-in slim list rows for schedules and approvals (#1703). - Batched lookups — session steering and citation lookups (#1704) and profile-sync banner lookups (#1701); agent status counts now run in SQL (#1716).
- Event-loop stalls stopped — KV range scans (#1690) and chunked app sync (#1696).
- SQLITE_BUSY — the final attempt is capped at 250ms with stall reporting (#1707); pricing and RBAC seed once per DB handle (#1695).
Runner and shutdown
- Graceful drain — the API drains on SIGTERM so workers hand off while it still serves (#1837).
- Stuck slots — worker slot freed when a provider session never settles (#1820).
- Output recovery — fenced JSON recovered for schema-bound tasks (#1833); pi reprompts once on an empty final turn (#1831); Claude task body sent once (#1815).
- MCP and workflows — unknown MCP session ids return 404 so clients reconnect (#1822);
node.retryhonored when an agent-task step fails (#1832).
Memory retrieval
- One relevance scale — every retrieval arm lands on one [0,1] scale with relevance-gated injection (#1689).
- Recall task content instead of worker wrappers (#1740); rater model pinned and recorded per rating (#1806).
- On main after v1.162.0 — stale chunks removed on keyed delete (#1857) and key-only lookup resolves one visible document (#1861).
Tasks and heartbeat
- Superseded tasks — dependents re-pointed to the resume (#1713, #1664); followUpConfig inheritance scoped to continuations (#1680).
- Heartbeat — flag parsing and interval 0 honored (#1711); fresh installs no longer point at unseeded scripts (#1714).
Breaking Changes
This week tightens behavior rather than removing APIs — these are requests that may now be refused:
- Task model validation — a task whose model the assignee harness cannot run is rejected at creation (#1764). In v1.159.0+.
- Task write ownership —
store-progressrefuses writes to another agent's task (#1737). In v1.159.0+. - RBAC, on main after v1.162.0 — stdio MCP servers at agent scope (#1851), changing a repo's allowMerge (#1852, #1859) and attributing a task to another user (#1853) now require a lead, operator or user. HTTP memory delete is gated and list is narrowed for agent callers (#1855).
Release Notes
Weekly release notes for Agent Swarm — new features, improvements, and fixes
Release Notes — Week of September 21 to September 28, 2026
Browser-first setup, catalog-only extensions that ship scripts and schedules, and per-credential usage plans, across 62 commits and 7 releases (v1.152.1 to v1.156.0).