Do code-graph tools actually save agent tokens? A 4-condition benchmark
Grep vs three code-graph tools on a real onboarding task: the effective ones halve the token bill — and one is worse than nothing.
Does a code-graph tool, used by an agent, save tokens / find faster / give a better answer than the same agent using only native grep? I benchmarked this on a realistic onboarding task (architecture, flows, spin-up, dependencies) against a mid-size codebase (~2k symbols), June 2026. Same frontier model across all conditions, one metered headless-agent subprocess per condition, fully separated contexts.
Conditions and headline numbers
| Condition | Tool | Turns | Wall | Billed input | Cost | vs grep |
|---|---|---|---|---|---|---|
| grep | native Read/Grep/Bash | 57 | 677s | 5.8M | $5.18 | 100% |
| codegraph | MCP graph server | 36 | 430s | 3.0M | $2.97 | 52% / 57% / 64% |
| poly | second graph MCP | 49 | 1167s | 4.9M | $5.74 | 84% / 111% / 172% |
| graphify | plain CLI | 46 | 501s | 3.2M | $3.14 | 55% / 61% / 74% |
Billed input = input + cache-creation + cache-read — the real token bill. Index builds are AST-only, 1–3s, no API tokens. Answer quality (blind LLM-judge, ground-truth checked): 9.50 to 8.50 out of 10 — within a point.
Findings
1. The two effective graph tools cut the token bill ~45% and wall-time ~30% at equal answer quality — the core thesis confirmed, for tools the agent actually uses.
2. The mechanism is fewer cache-writes, not fewer output tokens. Cache writes cost ~12.5× a cache read. Grep wrote 177k cache-creation tokens (kept reading files into context); codegraph wrote 122k, offloading reads to 4 structural queries. Graph tools replace file-reading: grep did 22 Reads, the CLI graph tool did zero.
3. Quality barely differentiates when docs exist. With a good CLAUDE.md present, every approach answers well — the graph tool's win is efficiency, not correctness.
A graph tool that's available but not leaned on is worse than grep.
The cautionary case proved it: one MCP condition's first run silently degenerated into a second grep (misconfigured server, zero tool calls). Even fixed, the agent used it twice and reverted to a 21-file Read loop — slowest (172%) and priciest (111%) of all four. From the token breakdown: it wrote 375k cache-creation tokens vs grep's 177k, because MCP tool definitions and discovery round-trips reshape the prompt prefix and invalidate the cache. It paid the MCP overhead and still did grep-style reading.
CLIs have a lower activation barrier than MCP tools for agents. Both MCP conditions initially under-used their tools — deferred tools require a discovery step the agent didn't know to take. An explicit "discover your tools first" instruction roughly doubled MCP tool usage.
A bonus axis: skill-authoring quality
Each condition also wrote a spin-up skill, validated by a fresh tool-agnostic agent on a busy machine. All four passed, but maturity differed: the condition that used a graph tool to truly understand the spin-up wrote the most robust skill (collision-safe ports, documented gotchas — 9/10); the grep condition assumed default ports were free (7/10).
Caveats: n=1 per condition — treat ~10% differences as noise, the ~2× gaps as real. Contexts fully separated; the harness was RED-verified before scaling; invalid first runs archived, not reported.