Runner Fleet Atlas
A self-hosted coding-agent fleet, reference-style: topology, lifecycle, identities, what's baked, the numbers, the footguns, and a ranked checklist for preparing agent-friendly environments.
One sentence: the vendor queues a session; your own always-on poller notices and launches a one-session VM from a pre-baked image, which registers with a single-use token, serves the session, then powers itself off into termination.
Nothing pushes into your network. The queue lives at the vendor; the only thing crossing the boundary is an outbound poll. Reference notes from operating one such fleet.
Three planes. The only inbound-shaped arrow is a poll the fleet itself initiates.
Components
| component | shape | job |
|---|---|---|
| orchestrator | small instance, 1 per environment, always on | poll the queue, stage the token, launch instances, keep one standby |
| session runner | 4 vCPU arm64, 100 GB root, 1 session | run the agent; terminate on exit |
| standby | identical launch, no session id | keep the environment selectable; absorb boot latency |
| launch template | per environment | pins the image, size, network, instance identity, tags |
| secret store | 1 environment secret per environment | read only by that environment's orchestrator |
| object store | durable scratch, stable prefix | state that must survive a resume |
| CDN-fronted bucket | separate from scratch | artifact URLs a human can open |
Environments
| default | long-running | |
|---|---|---|
| idle release | 90 min | disabled |
| wall-clock ceiling | disabled | disabled |
| startup timeout | 15 min | disabled |
| ends when | idle clock fires | a human archives it, or the box dies |
| warm standby | 1 | 1 |
| failure mode | parks mid-thought if it asks a question | bills indefinitely if abandoned |
- Same image, same code — the personalities differ by one number.
- Disabled clocks are written out explicitly, even where the module defaults them to zero: an inherited value is invisible at the call site, so nobody auditing the block against its own promise can see a knob that isn't there.
- Why each environment costs a permanently idle machine: the picker only lists environments with a runner checked in, and a dispatch to an environment showing no capacity generates no hint — so zero standbys means the environment never appears and can never bootstrap.
- The standby also hides boot latency. The ~226s is still paid, just afterwards by the replacement, off the critical path.
One session, start to finish. The conversation outlives the machine; the working tree does not.
What survives a resume
| survives | lives where | |
|---|---|---|
| conversation history | yes | vendor session stream |
| committed + pushed work | yes | your git host |
| uncommitted tree | saved, not restored | a rescue branch |
| background tasks, in-process teammates | no | — |
| the workspace | no | — |
- The vendor stores a pointer, your git host stores the content. The session record holds a branch name; resume fetches that ref from your remote.
- Proof from the failure mode: delete a branch after its PR merges and resume dies with a missing-ref error. Re-creating the ref fixes it — which it wouldn't if the conversation carried the code.
- Automatable: merged PR heads stay reachable forever on the forge, so the branch can be recreated server-side at its final commit with one API call, no clone.
- Resume often lands on a different machine, or on an existing idle standby — which may still be running the old image.
Identities
| identity | can do |
|---|---|
| orchestrator | launch instances (only ones tagged as fleet members), stage + delete registration tokens, read its own environment secret |
| instance | read the forge app key to mint short-lived tokens, archive session logs, fetch its own work order, pull container images |
| session | one object-store bucket, model invocation, one named secret. Nothing else. |
Four controls keep the session below its host — each alone proved insufficient
| control | covers | why the others don't |
|---|---|---|
| metadata service v2 + one-hop limit | containerised processes | the session isn't containerised |
| firewall rule rejecting the metadata endpoint for the session uid | the session user | sudo restores the path |
| blanket sudo stripped to narrow brokers | privilege escalation | found live, not theorised |
| shared-key environment sourcing removed | a key already in memory | makes the other three irrelevant |
Defence in depth fails when one layer hands out the thing the others protect.
The narrow-door pattern
- A session legitimately needs real app config; the tool that fetches it needs cloud credentials — the exact thing the firewall denies.
- Answer: one root broker per legitimate need, not a widened rule.
- Limits, all load-bearing: strict character allowlist on the argument, so a path can't escape the dev prefix and it can never return production; no passthrough flags, so it can't become an arbitrary invocation of the underlying tool; config only, never credentials.
- It pins its own cloud-config to a root-owned path — because that file is otherwise session-writable, and a credential-process entry in a config read by a root process is arbitrary root code execution.
The registration token
- Long-lived environment secret registers any runner and can claim any member's queued session — so it never reaches a host that runs session code. Orchestrator only, into a memory-backed filesystem.
- Executing host gets a single-use token, scoped to one session. Both go through the same CLI flag, which is the easy mistake.
- Staged in a parameter store, not instance user-data — user-data stays readable from inside the instance for its whole life.
- Encrypted at rest · written to tmpfs, never the snapshotted disk · deleted the moment it's read, so a later reader can't replay it · never logged.
- Lifetime matches the boot budget, not the session — a leaked staged token is inert within minutes.
Why runners cull themselves
- The permissions ceiling grants only the management agent's own registration calls. Remote command execution is absent, and no policy edit can add it — the ceiling caps what any role may grant.
- What is permitted: describe and terminate instances. So each runner decides for itself.
- Better anyway: a runner reads its own metrics, so it's never wrong about whether it is busy. Inferring another host's state from a launch-time tag can't distinguish a spare from a machine serving a person.
Baked vs ad-hoc
| baked into the image | ad-hoc per session | |
|---|---|---|
| repos | 11 full clones + warmed dependency install (~2 GB) | anything not on the list |
| toolchain | container runtime + build/compose plugins, forge CLI, k8s + IaC tooling, 3 language runtimes, language servers | — |
| agent config | operating contract, skills, subagents, hooks; environment fragment selected at boot | — |
| supervision units | yes — workstation-only ones deliberately disabled | — |
| databases / caches | no — client libraries only | containers, per test run, pulled cold |
| gitignored artifacts | no — generated schemas, env files, build stubs | reconstructed by hand — largest remaining per-session cost |
| secrets | never | fetched at boot into memory |
Which dependencies are real, which are containerised
| dependency | real or local | why |
|---|---|---|
| object store, secret store, parameter store, model endpoint, image registry | real | the session IAM identity reaches them directly |
| observability backend | real | traces are the verification surface |
| auth provider, feature-flag service | real | a fake bypasses exactly what a post-deploy check exists to exercise |
| tunnel provider | real | reserved wildcard domain + managed cert |
| forge | real | real org, real app installation |
| staging deployment | real | driven as a customer, asserted via traces |
| cache + relational DB | local containers | integration tests mutate and reset schemas; can't share with every other session |
| other OSS services | local build | no config service exists for them |
The line follows what the dependency is for: previews read the world as it is, so real services give a truer preview. Tests need isolation, so they get containers. The preview path never starts a database.
The toolkit
| skill | solves |
|---|---|
| preview | hand back an openable URL from a machine with zero ingress |
| smoke-test staging | prove a deploy works by driving it as a real user, then reading traces |
| verify units | tiered gates, with an explicit probe for whether a container runtime is even reachable |
| dashboard audit | the four independent layers in which a metric can lie |
| recycle runner | move onto a fresh image without waiting out the idle clock |
| subagent | question it answers |
|---|---|
| architecture review | is this the right shape? |
| quality bars | what would have to be true to call this done? |
| live end-to-end | does it work in the running system? |
| docs updater | what is now stale? |
| grill: industry practice | why is this not how the field does it? |
| grill: repo convention | why is this not how we do it? |
Each is a separate clean context. A reviewer sharing the author's context shares the author's blind spots.
| hook | event | effect |
|---|---|---|
| gap reporter | every tool call | turns environment defects into tracked issues |
| masked-failure guard | every tool call | flags success claims resting on a swallowed pipeline failure or truncated search |
| unpushed-work nudge | stop | blocks the turn when the checkout is dirty or unpushed |
| stale-image notice | stop | says the machine is running a superseded image |
Reaching a machine nothing can reach
- Zero ingress → no port to expose → previews go outbound through a tunnel on a reserved wildcard domain.
- Tunnelled HTTPS serves one port. There is no base URL with varying ports — every reachable service gets its own hostname.
- Static artifacts take a different path: a CDN-fronted bucket, deliberately separate from durable scratch, because a URL you hand someone shouldn't live in a bucket every session can overwrite.
The numbers
| quantity | value | note |
|---|---|---|
| cold boot to status-ok | ~226 s | dominated by init + service startup, not instance size |
| capacity re-evaluation | ~100 s | shorter than boot → one vacancy can yield two machines |
| spawn budget before the server retries | 300 s | must exceed real p99 boot; also how long a failed spawn goes unnoticed |
| idle release (default env) | 90 min | raised from 30 after self-paced loops kept dying |
| startup timeout (default env) | 15 min | re-arms on every attach, not just first init |
| git token lifetime | 1 h | minted root-side at boot; key confined to a subshell |
| registration token lifetime | minutes | matched to the boot window |
| drain budget | ~30 s | push the outcome branch before the workspace evaporates |
| standby cost | ~1 always-on instance per environment | the price of being selectable in the picker |
| abandoned no-clock session | ~3 days → a few dollars of nothing | nothing reclaims it |
Footguns, each observed live
| looks like | actually is |
|---|---|
| preview backend is broken — page loads, every API call fails | the frontend calls localhost; the reviewer's browser resolves that to their own laptop |
| tunnel plan limitation on a non-443 port | tunnelled HTTPS serves 443 only — there are no port-based URLs |
| wildcard cert not covering a host | the wildcard is one label deep — use a hyphen, not a dot |
| env var change had no effect | build-time inlined into the client bundle; restart required |
| appended override ignored | dotenv keeps the first occurrence of a key — replace in place |
| config fetch succeeded | a missing service returns empty output with exit 0 — test for empty, not exit status |
| all CI checks green | the checks output is TAB-separated; a \t pattern in ERE matches nothing |
| PR is unmergeable | mergeability is computed lazily — the first query returns unknown and triggers the computation |
| your own session died mid-cleanup | pkill -f matched your own shell's command line |
| runner never registered, no logs | console output is a differently-named permission than the broad describe grant |
| instance booted but the agent never started | a heredoc's trailing newline is load-bearing — cloud-init won't execute an unterminated final line |
| error legs missing from a rollup | lexicographic max() on a status column hides them — 's' > 'e' |
| metric oscillating like a sawtooth | day-bucketed query at sub-day resolution zero-fills between points; the peaks are the real values |
| dashboard panel shows “No data” | SQL returns hundreds of rows — the long-format series field doesn't render. Pivot to wide columns. |
| newest image by name | sort by creation date — re-bakes and gaps in a version sequence make name-ordering lie |
| a docker build failing on permissions | an inherited restrictive umask, set once at boot and inherited by everything downstream |
| a cloned repo missing its branches | a shallow clone pins the refspec to the default branch only |
| an unattended run that did nothing for 30 minutes | an unanswered permission prompt is a silent deny on a timer |
Control surface
| to change | mechanism | cost |
|---|---|---|
| idle / ceiling / startup clocks | module input → policy file on the orchestrator | apply only |
| image pin per environment | pin map, then roll the launch template + rotate standbys | apply + rollout |
| instance size, disk | module inputs | apply, with replacement |
| skills, hooks, subagents, baked repos | image build scripts | rebuild + roll |
| standby count, spawn budget | baked supervision unit | rebuild |
Audit periodically whether every knob you believe is applyable actually reaches the running process. A variable that is declared, documented, and never read looks identical to one that works.
An interactive agent asks. An unattended one loses the run.
Checklist: preparing an agent-friendly environment
Do first — hours of work, prevents lost runs
| prepare | because |
|---|---|
| pre-approve every tool permission the task needs | an unanswered prompt is a deny on a timer; the run dies having done nothing |
| block the turn on unpushed work | one stop hook; the machine is disposable by design |
| make shutdown terminate the instance | the exit path is the garbage collection — no cleanup daemon, no orphan accounting |
| write disabled knobs explicitly | an inherited default is invisible to whoever audits your config against its promise |
State
- Conversation on the platform, code on the forge, nothing durable on the box.
- A durable scratch store with a stable prefix across resumes.
- A separate artifact store for openable URLs.
- Rescue uncommitted work on exit by discovery, not declaration — a declared-checkout list misses the repo cloned on a whim.
Image
- Bake repos, dependency installs, toolchain. A stale bake costs a fetch; a missing one costs a cold clone every session.
- Bake the gitignored half too — generated schemas, env stubs, build outputs. Absent from a fresh clone by definition, and usually the biggest per-session cost.
- Warm service images if any test tier uses containers.
- Never bake secrets. Fetch at boot into memory.
- Pin per environment, so one can move while another holds.
Credentials
- The credential that can claim any work never touches a host that runs agent code.
- Executing hosts get single-use tokens, deleted on read, lifetime matched to boot.
- Narrow brokers, not blanket privilege — one command per need, each validating its input.
- Any config a privileged process reads must be owned by that privilege.
- Never a secret in a command line or a dumped environment — both land in the agent's own transcript.
Clocks
- Know exactly what defers your idle clock. “Defers while busy” and “defers while waiting to be busy” are different promises, and self-paced loops live in the gap.
- Every no-timer environment needs a reclamation path. At minimum, alarm on instance age.
- Distinguish a spare from a machine serving a person before anything may reap.
- Check whether spawn latency exceeds your capacity-evaluation interval.
Verification
- Assert on traces, not status codes: span exists, nested correctly, outcome says success, side effect landed, and nothing else changed.
- Read the warning level. A healthy response can hide four of five internal attempts failing.
- Two independent surfaces for any number the agent quotes. A mismatch is the finding, never noise.
- Assert a positive count. A job that matched zero tests is green, fast, and proves nothing.
What generalises
- Poll, don't listen. An outbound queue needs no inbound path → literally zero ingress → every other control gets easier.
- Make the machine the unit of work. One session per instance, self-terminating. No shared state to leak, no drift to accumulate.
- Split the credential by lifetime. Claim-anything secrets on the dispatcher; single-use tokens on the executor.
- Put state where it survives. Then losing a machine is an event, not an incident.
What doesn't generalise: the isolation model. Per-uid firewall rules on the metadata endpoint, narrowed sudo brokers, self-termination as the exit path — all shaped by a session running directly on a dedicated host. Move it into shared-kernel containers and none transfer; you'd re-earn each property with a different mechanism, and the failure mode would be silent.
The documentation trap
- The most expensive defect here was not code. A skill file claimed a lifetime guarantee that had been changed underneath it, so agents kept promising a ceiling that no longer existed.
- Another described a startup timeout as applying only before initialisation. It re-arms on every attach — that misreading cost an environment days of silent churn.
- Documentation an agent reads is executable, in the sense that the agent acts on it.
- So: keep measured constants out of prose nobody re-verifies, state the mechanism rather than the number, and tell the reader how to check the claim against the running system.