Do “token optimizer” tools survive contact with a real agent swarm?
Four community optimizers, nine configurations, three studies on an autonomous coding swarm: cost is run length, not verbosity — and only one combo reliably pays for itself.
Community “token optimizer” tools claim 60–90% savings — measured on one-shot CLI interactions. I evaluated four of them, plus combinations, on a real autonomous code-delivery swarm (a FastMCP control plane dispatching containerized Claude-Agent-SDK workers on an OSS coding model), across three studies in June–July 2026.
The contestants
| arm | configuration | mechanism |
|---|---|---|
| A | baseline | unmodified worker |
| B | caveman | terse-output persona |
| C | ponytail | anti-over-engineering persona |
| D | rtk | hook compressing bulky command output |
| E | rtk-opt | rtk + steering toward rtk-friendly commands |
| F | combo | ponytail + caveman |
| G | caveman-compressed | caveman + compressed operating contract |
| H | headroom | LLM-proxy context-compression callback |
| I | combo + headroom | F + H stacked |
Identical task prompt per arm; cost measured from wire-level spans (study 1) and the swarm's own repriced usage ledger (study 2), same rate table.
Results
Study 1 (frontend feature, n=3, wire-metered): combo+headroom −30%, combo −29% with the tightest variance of any arm (±$0.08), single personas −13…−18%, rtk +0.2% (dead even), rtk-opt +15% — worse than doing nothing. The single incorrect delivery of the whole program came from headroom's lossy compression.
Study 2 (backend deletion refactor, n=1): combo+headroom cheapest again with the fewest turns — but the middle of the field inverted: rtk-opt jumped to 2nd, ponytail fell to last. Single samples cannot rank the middle: within-arm spread across repeated runs ($3.23–$5.52) is wider than most between-arm gaps.
Why: the money is in re-sent context
Cache-read of the re-sent conversation is ~⅔ of every run's dollar cost in both studies. The arms that win finish in fewer turns — each avoided turn removes a full re-send of the multi-million-token cached context. Trimming per-call output attacks a small slice and does not move the total. That's also why an input-layer command compressor doesn't transfer: the agent emits compound multi-command scripts that bypass the rewriter (~6% hit rate even under steering), and the steering itself converted batched scripts into extra round-trips — how rtk-opt backfired into the most expensive arm.
Prove the mechanism fires in the exact runtime path you're measuring — an integration verified on the CLI path proves nothing about the SDK path.
The integration trap (the finding I'd keep if I could keep one)
The personas inject via Claude Code hooks — and hooks configured in settings.json do not fire in an SDK query() worker. They fire only in the interactive CLI. A live worker with the persona “enabled” produced a transcript with zero persona activation lines: the integration shipped as a silent no-op. Moving injection to in-process SDK options.hooks (fired at run start and on every re-drive, plus subagent injection) is what actually works — verified by transcript, and the combo then saved −23% tokens / −24% cost on clean production hardware.
Methodology scars (all survivable, all instructive)
- Base contamination: the study-2 task was anchored to an open PR that merged mid-experiment; later rounds cloned a base where the task was already done and every arm degenerated to a cheap near-noop. Cost was low because there was nothing to do. Freeze the base or pin the clone commit.
- Structural arm identity beats shared state: a shared image tag got clobbered by a concurrent session; immutable per-arm images and per-arm queues fixed it, plus a detached driver surviving session teardown.
- Don't trust the client's price map: the CLI booked the OSS model at frontier rates ($9.61 reported vs $3.26 ledger-truth). Reprice from wire tokens against a pinned rate table.
Verdict
Ship the persona combo: −29% at n=3 with the tightest variance and zero correctness incidents, no model-path infrastructure. Skip input-command compression for cache-dominated agents. Judge every future optimization by one question: does it shorten the run?