A traffic light for AI agents sharing one Mac
Netdata alerts become a GREEN/YELLOW/RED verdict that every Claude Code session reads before heavy work — advisory, actionable-only, with the fix in the warning. What it watches, how agents hear it, and what a six-reviewer panel changed.
Many AI agents share one Mac. None of them can tell when the machine is struggling, so they keep starting builds, containers and new agents until everything slows down. The fix is a traffic light agents can read: GREEN go, YELLOW finish what you are doing, RED stop. It only advises — nothing is blocked — and agents follow it because the warning lands in their context together with the exact fix to apply.
An alert stays on only if an agent can act on it, and its text names that fix.
Signals flow up into Netdata alerts; the worst raised alert becomes the verdict every session reads.
The alerts
Netdata ships with about 180 stock alerts. All but a handful are switched off. What remains:
| Alert | Fires when | Fix it names |
|---|---|---|
| CPU | above 85% for 10 min | kill the listed runaway |
| RAM | app + wired + compressed memory above 85% | close idle sessions, stop containers |
| Swap-in | reading swap back from disk above 10 MiB/s | same — memory has run out |
| macOS pressure | the kernel itself reports critical (level 4) | stop starting work, free memory |
| Runaway process | one of your processes holds a full core for 10+ min | safe kill of that PID |
| Docker VM memory | the VM above 80% of its RAM | stop the biggest unused container |
| Container restart loop | a restart in the last 5 min | read its logs, fix or remove |
| Disk | nearly full, or full within hours at the current rate | prune Docker, clear caches |
Deliberately not alerted
- Swap used. macOS frees swap lazily and leaves it high for hours after the pressure is gone, so an alert on it never clears. The read-back rate is what separates a sick Mac from a healthy one.
- macOS pressure level 2 (warn). This Mac sits there on a normal afternoon at 35–45% free with nothing to fix.
How agents hear it
| When | What happens |
|---|---|
| You send a prompt | If not GREEN, the verdict, reasons and fixes are added to the agent's context. |
| An agent runs a heavy command (docker run, a build, a test suite, a new agent) | The same warning arrives with the command's result. |
| An agent checks before heavy work | A small CLI exits 0 GREEN, 1 YELLOW, 2 RED, 3 UNKNOWN (monitor not updating). |
Interactive sessions treat YELLOW as finish current work, start nothing heavy. Autonomous runs cannot wait, so for them the verdict changes how they proceed, never whether they end: YELLOW means one heavy step at a time with no fan-out, RED means write an escalation and stop, and a run never kills a process it did not start.
Recovery is automatic
Nothing latches. The verdict is recomputed every 10 seconds, so it returns to GREEN as soon as the alerts clear. Alerts clear only after the value drops below a lower bar and stays there for a while (15 minutes for CPU, RAM and swap), so the light follows sustained pressure, not spikes.
Guardrails
- Kills go through a checker. A fix line says
mac-pressure.py --kill PID, not a barekill. It refuses unless a fresh verdict still lists that PID with the same start time and executable — a PID quoted several turns ago may belong to a different process by now. - No command lines reach agents. Runaway rows carry only the executable name. A full argv can hold a bearer token, or text posing as instructions, and this text is injected into every session.
- The hook cannot break work. It never blocks a tool call, and if its script is missing it does nothing.
- A dead monitor is not "all clear". If the verdict file stops updating, agents are told UNKNOWN rather than nothing.
What a review panel changed
Six fresh-context reviewers — architecture, simplicity, security, test-hardening, performance, blast-radius — went over the first version. The findings that mattered:
- The light was always on. The plugin judged the kernel's pressure level itself, outside the alert rules, and level 2 is normal here. Now the level is charted and alerted like everything else, critical only, so every threshold lives in one place.
- The monitor fed the pressure it reported. It listed sessions every 10 s by launching a 161 MiB Node CLI. Now once a minute.
- The tests mostly could not fail. 23 of 27 deliberate mutants survived, including one where the verdict was always GREEN. After rewriting the checks against fakes, 26 of 26 die, each on the case that names it.
- One disk alert could never fire. It read a helper value the whitelist had switched off, so it was permanently null.
- Mac alerts were also running on the Docker VM, with Mac fix text. Now scoped by host label.
Limits
- It is advice; an agent can ignore it.
- Only active sessions hear it. The idle sessions that actually hold the memory never get a prompt, so a person still has to close them.
- A warning on a heavy command arrives after the command has started.
Takeaways
- Give agents a signal and a fix, not a number.
- Delete every alert nobody can act on; noise teaches agents and people to ignore the light.
- Pick signals that measure pain (swap read-back), not ones that merely sit high (swap used).
- A light that is on most of the time is off. Measure how often it is lit before trusting it.
- Start advisory; enforce only after the signal has earned trust.