Human off the loop — notes on autonomous coding agents
Interactive and autonomous are two lanes, not one. What it takes to let a coding agent's loop close itself.
I spent the last few months building an autonomous coding-agent platform. This is the technical recap — what the system is, what actually works, and what's next.
The one-line thesis: interactive and autonomous are two lanes, not one. Teams that release to production several times a day do it with a mechanism, not talent — and the mechanism is a loop the agent can close itself.
Three eras of using AI to code
| Era | Mode | What it looks like |
|---|---|---|
| 2021–24 | Human in the loop | Completion, then chat that edits files. Human runs the tests, reads every turn. Human is the loop. |
| 2025– | Human on the loop | Agent in the shell with real tools, repo rules, subagents. Plan → act → verify → fix → ship. Runs unattended; human still merges. |
| next | Human off the loop | Agent owns the goal, not the ticket. Picks the work, ships it, operates it. Human sets direction — nothing else. |
The efficient way to use AI is to set the goals and the test checks up front, then let the agent loop inside them on its own. That takes back the time spent waiting at the keyboard hitting continue, and the time spent reading AI code line by line. Without it, what AI gives you stays capped by human energy and hours.
Four things that work
Loops — a feedback loop the agent can act on. Every CI event re-drives the run. PR review comments get resolved automatically. No human turn in between. The loop closes itself.
Guardrails — it knows when to stop. Budget-aware, capped per run. Re-drive caps and a dead-letter path on repeated failure. A run ends when the goal is met, not when the tokens run out.
Models — a quarter of the frontier price. Multi-model routing puts an OSS model on the coding work at ~25% of the frontier-model cost, with skills and orchestration holding quality. OSS-capable is not a downgrade.
Control — a state machine, not a black box. Resume a run by MCP tool call or by PR comment. Status reported back while it runs. A chat notification when it's done. Steerable from outside.
Two lanes
| Interactive | Autonomous | |
|---|---|---|
| Who drives | Human drives, agent answers | Agent drives, agents audit, human signs off |
| Input | Vague, taste required | Bounded, against a clear spec |
| Verification | Human reads every turn | A verifier calls it done |
| Ships | A decision, a spec | Merged code |
Complementary, not competing. A normal day uses both: brainstorm in a session → dispatch the work → close the laptop → a PR comes back with suites green, multi-agent reviewed, and a running app URL to click.
The run loop: build → CI events → dev env → live e2e → review bots → fix. Runs until green, then posts the link.
What a run looks like
| # | Stage | What happens |
|---|---|---|
| 1 | Build | Agent writes on its own branch. |
| 2 | CI events | Suites run. Results stream back. |
| 3 | Dev env | Standard, claimable environment — same shape for every service. |
| 4 | Live e2e | Tests drive the deployed app. |
| 5 | Review bots | Static checks. Review rounds. Comments resolved in-loop. |
| 6 | Fix | Any red re-drives the run. |
What's next
- Environments — a standard, claimable dev env for every service, so the agent can run e2e across systems and end on a running URL a human can click. Unblocks everything else.
- Feedback quality — autonomy only opens the gate; the signal matters as much as the loop. Tests broken down far enough to point at the fault, and multi-agent adversarial review on the diff instead of one pass. A loop fed a weak signal just fails faster.