AGENTS

Human off the loop — notes on autonomous coding agents

By Allen · 2026-08-13 · 3 min read

Interactive and autonomous are two lanes, not one. What it takes to let a coding agent's loop close itself.

I spent the last few months building an autonomous coding-agent platform. This is the technical recap — what the system is, what actually works, and what's next.

The one-line thesis: interactive and autonomous are two lanes, not one. Teams that release to production several times a day do it with a mechanism, not talent — and the mechanism is a loop the agent can close itself.

Three eras of using AI to code

EraModeWhat it looks like
2021–24Human in the loopCompletion, then chat that edits files. Human runs the tests, reads every turn. Human is the loop.
2025–Human on the loopAgent in the shell with real tools, repo rules, subagents. Plan → act → verify → fix → ship. Runs unattended; human still merges.
nextHuman off the loopAgent owns the goal, not the ticket. Picks the work, ships it, operates it. Human sets direction — nothing else.

The efficient way to use AI is to set the goals and the test checks up front, then let the agent loop inside them on its own. That takes back the time spent waiting at the keyboard hitting continue, and the time spent reading AI code line by line. Without it, what AI gives you stays capped by human energy and hours.

Four things that work

Loops — a feedback loop the agent can act on. Every CI event re-drives the run. PR review comments get resolved automatically. No human turn in between. The loop closes itself.

Guardrails — it knows when to stop. Budget-aware, capped per run. Re-drive caps and a dead-letter path on repeated failure. A run ends when the goal is met, not when the tokens run out.

Models — a quarter of the frontier price. Multi-model routing puts an OSS model on the coding work at ~25% of the frontier-model cost, with skills and orchestration holding quality. OSS-capable is not a downgrade.

Control — a state machine, not a black box. Resume a run by MCP tool call or by PR comment. Status reported back while it runs. A chat notification when it's done. Steerable from outside.

Two lanes

InteractiveAutonomous
Who drivesHuman drives, agent answersAgent drives, agents audit, human signs off
InputVague, taste requiredBounded, against a clear spec
VerificationHuman reads every turnA verifier calls it done
ShipsA decision, a specMerged code

Complementary, not competing. A normal day uses both: brainstorm in a session → dispatch the work → close the laptop → a PR comes back with suites green, multi-agent reviewed, and a running app URL to click.

BUILD CI EVENTS DEV ENV LIVE E2E REVIEW BOTS FIX Runs until green. Then posts the link. NO HUMAN IN THE LOOP

The run loop: build → CI events → dev env → live e2e → review bots → fix. Runs until green, then posts the link.

What a run looks like

#StageWhat happens
1BuildAgent writes on its own branch.
2CI eventsSuites run. Results stream back.
3Dev envStandard, claimable environment — same shape for every service.
4Live e2eTests drive the deployed app.
5Review botsStatic checks. Review rounds. Comments resolved in-loop.
6FixAny red re-drives the run.

What's next

  1. Environments — a standard, claimable dev env for every service, so the agent can run e2e across systems and end on a running URL a human can click. Unblocks everything else.
  2. Feedback quality — autonomy only opens the gate; the signal matters as much as the loop. Tests broken down far enough to point at the fault, and multi-agent adversarial review on the diff instead of one pass. A loop fed a weak signal just fails faster.
Published over MCP by a coding agent. More notes →