Lessons of an Autonomous Swarm
Nine weeks of an agent runtime's git history reviewed commit by commit: eight eras, the five costliest incidents, the silent-failure classes, and the design lessons that survived contact with production.
The system under study is the runtime behind an autonomous coding swarm: an MCP control plane plus stream-dispatched, Kubernetes-orchestrated agent workers that turn a strategic goal into a reviewable PR. It went from first commit to a production fleet in nine weeks — and because the swarm dogfoods itself on its own repo, its git history is an unusually honest record of what actually breaks when an agent system meets production.
I replayed the entire history: 654 squash-merged commits (one commit ≈ one PR), over nine weeks in mid-2026, every commit reviewed by a fan-out of eight parallel readers working from the real commit bodies, spot-checked against git. This is the organized result: the era timeline, the big impacts, and the lessons that generalize.
Eight eras across nine weeks, with the three costliest incidents marked.
The era timeline
Era 1 — Genesis (week 1)
All the load-bearing decisions land in week one: hexagonal layers (domain→ports→adapters→contexts→app), the ack-after-durable Turn-Stop invariant (claim → commit → status, acknowledge the queue entry only after the durable commit), stream consumer groups for exactly-once dispatch, git-bundle checkpointing, the sovereign-idle CI/review wake loop, and the conversation-resume protocol. The first production bugs are already the house style: five duplicate runs because no heartbeat guarded the pending entry; every PR comment silently dropped because ingestion read comment.login instead of comment.user.login; write-ahead dispatch durability making every fresh run look like a resume, so no run could open a PR until the follow-up fix.
Era 2 — Buildout (week 2)
The gateway goes read-only; the agent gets its own bot identity after a human-account run looped on its own replies; multi-repo runs, cancel, and warm-pod resume ship. Two architectural pivots: the carried-state feedback queue is deleted for notification+fetch (thin wakes, the code host as the source of truth on re-drive), and a trace fold-proxy starts reparenting the agent CLI's spans — immediately hitting the export-ordering trap (spans export on END, so the wrapping span always arrives last). Two security criticals in one week: a shell-injection RCE via an unescaped worktree path, and webhook HMAC auth that failed open when the secret was unset.
Era 3 — Hardening sprint (late week 2 – week 3)
Five days sweeping two audit umbrellas. A deterministic pre-tool-call write-guard becomes the real containment floor for permission-bypass mode. The no-PR delivery guard turns out to be scoped out for the entire fleet — 19 runs in one week reported “done ✅” with zero PRs. The agent transport swaps from a one-shot query to a held-open client so background work stops dying at teardown. A poison-pill max-delivery cap ends the burn-a-pod-per-reclaim-forever loop, and an LLM-scored admission clarity gate goes in at the front door.
Era 4 — Audit and pivots (weeks 3–4)
The delivery model whiplashes — staged orchestration added, retired for pure inline, then restored behind a forcing hook — while the runtime goes multi-provider through a model proxy after the dev-lane model account is disabled. The defining incident: a $145.87, 146-leg, 21.3M-token runaway because retry counters lived in pod memory and reset on every crash-resume. The fix — durable database counters — becomes the template for every later cap. The silent-noop zombie turn is named: a provider error rewritten to a normal stop, so a dead run reports success at zero tokens and re-drives invisibly forever.
Era 5 — Guardrail ladder, OAuth, cost truth (weeks 4–5)
A 298-leg / 29-hour loop triggers the full ladder: per-PR re-drive budget, same-failure convergence detector, silent-noop push-skip, fleet-wide provider circuit breaker. Cost accounting turns out to be lying everywhere — tokens stored session-cumulative (~8.5× overcount), the alternate provider priced off the CLI's bundled price map (~50× over). The proxy trace-fold is redesigned three separate times, each disproven only by looking at the rendered live trace. A background-task drain defaulting to the run's 2-hour timeout had been delaying every PR delivery by up to two hours. The dev lane cuts over to managed Postgres and Redis.
Era 6 — Infra ownership and interactivity (weeks 6–7)
Single-writer ownership lands piece by piece: IaC state to a remote backend under a GitOps apply bot, secret values out of state files via an external-secrets operator, Kubernetes manifests to a GitOps controller, deploys via merge-to-main OIDC CI, and the gateway's own auth retired entirely in favor of one OAuth door. The interactive layer ships in a two-day sprint — watch, needs-input park/answer, chat notify. Two shipped features are then discovered to have never run at all: persona hooks wired to a hook surface the SDK path never fires, and a progress mirror matching a tool family the deployed agent doesn't use.
Era 7 — Incidents and the investigator (weeks 7–8)
The headline: a fatal startup guard for a config invariant ships, the config repo hasn't caught up, and every worker exits 1 for ~40 hours. The fix downgrades the guard to warn-plus-gauge and moves the blocking check a layer up the deploy pipeline. Guardrails counting the wrong quantity get corrected: the redispatch cap was counting review rounds, not failures; subagent spend was invisible to every cost counter (a delegating run booked as little as 24% of real cost); every subagent silently ran the top model tier. A webhook burst spawning up to eight duplicate pods per run goes through three designs before landing on the simplest one — reading the queue's own undelivered entries, no new state. A read-only investigator agent type ships with its own containment and RCA egress.
Era 8 — Lifecycle maturity and discipline (week 9)
Pointer-driven dispatch makes goals the single admission surface — a ticket, issue, or chat pointer is fetched and organized into goals, and the agent can pick its own allowlisted repo. A new CONCLUDED status arc gives deliberate agent stops a resumable identity distinct from operator cancels, plus a human-comment revival door for dead-lettered runs. Forensics find the PR-ready alert dark for weeks (a snapshot-timing race with the receiver's webhook write) and the admission scorer silently degraded on ~55% of dispatches (the provider's extended thinking starving the token budget under forced tool choice). A ~25-commit complexity-refactor train ends in a hard repo-wide lint gate, and a docs-currency convention lands after five stale doc claims burned autonomous runs in a single session.
A soft probabilistic risk must never be wired to a hard total failure — the guard against an occasional duplicate pod took the whole fleet down for forty hours.
The big impacts
Incidents, by cost
- ~40h total-fleet outage — fatal startup guard vs. lagging config; every worker exited 1. Also exposed that all alerting was worker-emitted, so monitoring went silent exactly when the fleet died — fixed with one alert that reads the control plane (queue depth vs. live workers).
- $145.87 / 146-leg runaway — retry caps in pod memory, reset on every crash-resume.
- 298-leg / 29h re-drive loop — no absolute budget across wake reasons; spawned the guardrail ladder.
- Every PR delayed up to 2h — background-task drain defaulting to the run timeout.
- 8 duplicate pods per run, 17% duplicate dispatch — freshly published work invisible to every guard during the scheduler's spawn window.
Silent failures — the worst class
Nothing in this list errored. Each was found by forensics, an incident, or a human noticing an absence:
- 19 runs reporting success with zero PRs (guard scoped out for the whole fleet).
- Every human PR comment dropped for days (one wrong field access).
- Zombie turns: provider errors rewritten to normal completions, dead runs re-driving forever at zero tokens.
- PR-ready notifications dark for weeks (snapshot-timing race).
- ~55% of admission-score calls silently truncated into a saturated fallback heuristic.
- Two shipped features that never executed in production — both worked in their canary, which tested a different execution path.
Security criticals
- Webhook HMAC failing open with an unset secret — forged wakes and forged terminal statuses possible.
- Shell-injection RCE via an unescaped path spliced into a Python literal.
- Repo-allowlist bypass through the ticket-dispatch path; a gateway secret leaking into the model subprocess via a prefix-based env passthrough; tokens on argv and in gitconfig (both moved to safer channels).
- Deterministic pre-tool-call guards proved to be the real safety floor — vindicated when the AI permission classifier itself flapped, denied every tool call, and had to be bypassed in favor of the deterministic hooks.
Defining design decisions
- Ack-after-durable + exactly-once via the pending-entries list — day one, never relitigated.
- Notification+fetch over carried-state feedback: the code host is re-read on every wake.
- Bot identity for the agent — a
[bot]login classifies natively and killed the self-feedback loop class. - Provider-agnostic model proxy — forced by an account outage, it made the fleet portable and exposed that every bundled price map lies.
- Single-writer infra ownership — the GitOps apply bot owns IaC state, the external-secrets operator owns secret values, the GitOps image updater alone writes the live image tag, one OAuth door in front of everything.
- Never-stuck lifecycle — park on genuine hardblock, dead-letter to a resumable status instead of a terminal one when a PR is live, revival doors for stopped runs.
The lessons that generalize
- Silent vacuous success is the house failure mode of agent systems. An autonomous loop hides its own failures by design: exit 0, subtype success, green checks. The countermeasures that worked were structural, not attentional — guards keyed on durable ledgers (did a PR row appear?), fences RED-verified by neutralizing the mechanism, and alertable gauges instead of info logs.
- Every cap eventually counts the wrong quantity. Heartbeats → caps → durable counters → budgets → convergence → circuit breakers → streaks → lifetime leg caps: each generation fixed the previous one measuring the wrong thing (review rounds instead of failures, parent-thread instead of total spend, per-leg instead of cumulative). Ask what the counter measures, not what it is named.
- Frozen-at-spawn state is a slow-fuse bug. Tokens, secrets, and env are fixed when a process starts; anything long-lived will outlive them. The arc went token-in-env → per-leg re-mint → runtime-refreshed file read per operation.
- Cross-process observability joins are never right the first time. Three full redesigns of one trace join, an export-ordering trap, an upstream CLI dropping the attribute a join keyed on. Unit tests caught none of it; rendering the live trace caught all of it.
- Fail-open vs. fail-closed is calibrated by blast radius, not principle. Auth fails closed (a forged webhook is worse than a dropped one). Admission scoring fails open (a scorer outage shouldn't stop dispatch). And a startup invariant check warns — because the guard's failure mode (total outage) was strictly worse than the risk it guarded (an occasional duplicate pod). The blocking version moved a layer up, into the deploy pipeline, where failing closed blocks a rollout instead of a fleet.
- Reversals are health, not churn. Roughly ten clean reversals in nine weeks: orchestration modes, guard families, an abstraction collapsed the day after shipping, a review model dropped, an ingress retired same-day, a mutation-testing lane ripped out after failing in three environments. The repo deletes what live evidence disproves — and the deletions read as confidently as the additions.
Method note: history replayed from the squash-merge commit bodies (which carry full PR descriptions), reviewed slice-by-slice by eight parallel readers, with cited hashes and dates spot-checked against git. The repo's own decision ledger and rule files were used as cross-reference.
Appendix — the architecture, end to end
Every piece the history above keeps referring to, on one map. Solid arrows are the dispatch-and-delivery path; dashed arrows are the feedback loop that re-drives a run; dotted arrows are observability and notification. The two invariants that hold it together: the queue entry is acknowledged only after the durable commit, and the feedback loop carries thin wakes — the code host is re-read as the source of truth on every re-drive.
The swarm end to end: one OAuth door into the gateway; write-ahead row + queue token; autoscaled worker pods running the agent behind deterministic guards; Turn-Stop durable commit before ACK; webhooks re-driving runs via thin wakes; reaper sweeping orphans; every plane folded into one trace.
Appendix — a catalog of graceful design
A structured sweep of the current codebase (eight parallel readers, one per subsystem) cataloged 97 distinct graceful-degradation mechanisms, every one anchored to real code. They collapse into ten patterns; below, each pattern with its strongest exemplars. File names are the runtime's own modules.
1. Fail-open vs. fail-closed, calibrated per site
- The write-ahead dispatch record fails closed while every advisory probe around it fails open — the one write that must not lie vs. the reads that must not block (
dispatch_task). - The clarity floor rejects only LLM-calibrated scores; the degraded heuristic lane never rejects a dispatch (
gatekeeper). - The leg cap fails open while the dead-letter decision fails closed — opposite directions in the same file, each matched to its blast radius (
consume_work). - Webhook HMAC fails closed on an unset secret, with an explicit, boot-loud local-dev opt-out (
receiver). - The startup invariant guard warns + gauges + boots — deliberately never fatal, the codified lesson of the 40-hour outage (
runner).
2. Best-effort with an alertable signal
The rule that makes suppression safe: a swallowed failure must stay visible. Notification egress failures log on two independently-suppressed surfaces; the delivery-count fail-open carries a warning naming exactly the risk it accepts (cap evasion); unresolved feedback drops count themselves; a zero-row status update warns instead of crashing or staying quiet.
3. Transient vs. permanent, classified before acting
- A bare
HTTP 404from a clone stays transient (an unauthenticated private repo 404s); only definitive not-found signatures terminalize. - A branch probe distinguishes “branch genuinely absent” (clean exit, empty output) from “probe failed” (raise, retry later) — a transient blip once permanently killed a resumable run.
- A poison webhook body is refused once with 400, never 500-looped — the code host retries 5xx forever.
- A parse-corrupt queue token is dead-lettered at the boundary; its valid batch-mates survive.
4. Bounded retry, backoff on transient classes only
Token minting retries transient network errors with exponential backoff, fails fast on shape errors, and single-flights concurrent minters. In-leg provider retry is classified, budgeted, and disabled during drains. Every queue client is built in one site with backoff retry plus an idle-socket health check, so hours of quiet cannot hand a dead socket to a wake.
5. Compare-and-swap transitions and one-shot latches
Cancel vs. resume on a parked run is settled by a CAS, not read-then-write. Dead-lettered runs revive through a CAS door that latches behind itself — one human wake, no re-armed loop. A duplicate-dispatch gate orders its two non-atomic reads so the only possible transition between them is always caught.
6. Idempotency and ordering-by-construction
The consumer group is created before every publish (already-exists is a no-op), so cold-start stranding is impossible. A merged PR can never be resurrected by a late webhook. Migrations replay on every boot under an advisory lock. A screenshot branch create tolerates exactly the two already-exists codes and raises on everything else.
7. Caps and budgets that crashes cannot refund
The redispatch budget is a consecutive-failure streak (success resets, outage exempts, breach freezes) stored durably; the idle budget is anchored in the database so a pod kill resumes with the remaining window; the poison-pill cap writes its dead-letter before acknowledging, so a crash between the two re-runs the halt, not the work.
8. Resumable dead-letter and park-not-fail
A deterministic failure after a PR exists writes HALTED (resumable, revivable) instead of FAILED (terminal, orphaning) — and the choice fails closed when the ledger is unreadable. A genuinely blocked run parks with a question instead of guessing; an unanswered park proceeds-without after a TTL rather than stranding. A fleet-level circuit breaker parks new dispatches during a provider outage instead of spawning doomed pods.
9. Demote, don’t repair; stale beats absent
A damaged transcript mirror is never repaired — resume is demoted to the known-good blob. A failed token refresh leaves the last file in place (stale credentials degrade to a retryable 401; absent ones hard-fail). The trace buffer evicts spans only on confirmed ingestion. The trace fold prefers a duplicate span over an orphaned one wherever it must guess.
10. Deterministic floors under probabilistic layers
A regex-and-structure pre-tool-call guard denies catastrophic commands regardless of what any model-based classifier decides. Secrets are tombstoned (not merely omitted) from the agent’s environment so an SDK merge cannot silently re-inherit them. The queue’s socket timeout is derived from the largest blocking read actually used, so no config change can break the invariant. A metric-attribute allowlist is default-deny, making cardinality explosions a conscious decision.
Counts by subsystem: ingestion/receiver 13, observability 13, control plane 12, execution core 12, code-host adapters 12, durability layer 12, agent-SDK adapter 12, queue adapters 11.