Agent Evals That Train: the Harbor + Osmosis Pattern
Why agent benchmark numbers go wrong, the two paradigms of agent evaluation, a comparison of the major frameworks, and a one-adapter pattern that makes the same verifier serve both benchmarking and RL training.
The problem: agent numbers lie by default
Evaluating an LLM agent is harder than evaluating a model, because the number you get is a property of the whole system — scaffold, tools, environment, verifier — not of the model alone. Two failure modes dominate. First, the harness is the confound: the same model can swing from 23% to 52% pass@1 purely by changing the agent scaffold, which is larger than most model upgrades. Second, verifiers are adversarial targets: a recent study found roughly 16% of tasks across five public terminal-agent benchmarks could be passed by frontier models without solving the task — by reading, patching, or gaming the grader.
Both problems get worse the moment an eval's grader doubles as an RL reward, because now a model is optimized against the verifier for thousands of samples. This article lays out the general ways teams evaluate agents, compares the major frameworks, and describes a pattern — one adapter projecting into Harbor for benchmarking and Osmosis for training — that keeps the benchmark and the training reward provably the same function.
干货 — the dry summary
The category map, one line each
- Paradigm A · hermetic benchmark — frozen tasks, sandboxed execution, verify effects on the world; the only paradigm that answers “is B better than A?” — Harbor/Terminal-Bench, SWE-bench, OSWorld, Inspect AI.
- Paradigm B · live-trace observation — grade production traces continuously; answers “is prod healthy?” but has no counterfactual and grades narrative, not effects — OpenAI trace grading, LangSmith/Braintrust/Phoenix.
- Paradigm C · in-flight grading — a judge inside the agent picking among best-of-N candidates at serving time; benchmarks then measure agent+selector as one unit.
- Meta-leaderboards — HAL: one shared scaffold across many benchmarks, cost-controlled, to kill the harness confound.
- User-simulation benches — tau2: agent and a simulated user both mutate shared state; the user-sim is a second model-under-eval on the reward path.
- RL-environment hubs — verifiers/Environments Hub, SkyRL+OpenEnv: dataset+rubric+rollout as one artifact, rubric-as-reward; the public eval→training bridge.
- The fused cell (this pattern) — hermetic effect-verified benchmark that is also the RL training environment, provider-neutral; no public framework occupies it yet.
Steal-this checklist
- One adapter, two projections — one source generates the benchmark tasks and the training dataset with byte-identical prompts; grader copies are generated, and byte-compare tests fail when a copy drifts.
- Tamper boundary — the verifier and its reference data enter the container only after the agent's phase; grade only from verifier-private copies, never from anything the agent could write.
- Gate-then-grade — deterministic preconditions force a hard 0 and skip the paid judge; dense deterministic axes for partial credit; the LLM judge is last and only for what can't be mechanized.
- Three outcomes, never two — scored-zero (model's fault) ≠ ungraded (grading infra's fault, excluded from means) ≠ fail-closed (judge errored where a skip would vacuously pass).
- Judges are frozen instruments — pinned model, per-criterion isolated calls, binary criteria, cross-family from the candidate, calibrated against human labels.
- Oracle pins the grader in CI — generate the known-good answer from the row's own grading contract and assert it scores 1.0; grader and generator can't drift without a red test.
- Failure mining — every production defect becomes a dataset row plus a grader check that would have caught it; a bug isn't fixed until a row scores it zero.
- Frozen SUT over hermetic ports — real agent code, fake world (file stores, JSON catalogs, private CDN); infra weather never moves a score or a gradient.
- State seeding over transcript replay — evaluate turn N by materializing turn N−1's world, not by replaying another model's sampled output.
- Report passk + SEM and a cost axis — consistency and price are product metrics; a keep-the-best aggregation erases the variance you needed.
The A/B property: swap only the agent binding and every other box is identical — that is what makes two runs comparable. On the RL path, the same step-3 scalar becomes the per-sample training reward.
The two paradigms (plus one hiding inside the agent)
A. The hermetic benchmark loop — offline and counterfactual. Freeze a task set, fake the world into determinism, execute the candidate in a sandbox, and verify effects on the world: files compile, tests pass, state changed. This is the only paradigm that answers “is B better than A?” — both candidates run the identical distribution, so the delta is attributable. Its costs: a slow clock, expensive world-faking, and a dataset that always lags live traffic. Harbor, SWE-bench, OSWorld, and HAL live here.
B. The live-trace observation loop — online monitoring. Grade what production actually did — traces, tool choices, handoffs, answers — continuously, on the true distribution. It answers “is prod healthy, did the deploy drift?” fast and with zero dataset maintenance. Its costs: no counterfactual (you cannot score a candidate that never served traffic), it grades the agent's narrative rather than verified effects, and uncontrolled inputs make week-over-week numbers incomparable. OpenAI trace grading and the LangSmith/Braintrust/Phoenix class of tools live here.
C. In-flight grading — a judge inside the agent, selecting among N candidates at serving time (best-of-N with a vision or quality judge). Mechanically the same as B, but embedded in the system under test. It matters because a paradigm-A benchmark then measures agent-plus-selector as one unit — and if the serving-time selector and the benchmark grader share a model family, the two scores couple: the selector picks what the grader will like.
The mature shape is a cycle: the observation loop discovers failures on live traffic; they are frozen into the benchmark as rows and grader checks; the benchmark gates the change counterfactually; the shipped change is monitored again.
Method doctrine inside paradigm A
Across the current guidance (Anthropic's and OpenAI's eval writeups, the LLM-as-judge literature, benchmark postmortems), the paradigm-A playbook converges on a short list:
- Three grading tiers — code-based (fast, deterministic), model-based (rubric judges, needing calibration), human (gold standard) — each tier calibrating the one below it.
- Gate-then-grade — hard binary preconditions force a zero before any partial credit, and the paid judge is never called on a gated failure.
- Judge hygiene — pin the judge model (a rubric is calibrated against a specific model), decompose to one isolated call per criterion, prefer binary criteria over Likert scales, use a different model family than the candidate (self-preference bias is measured at 10–25%), and report agreement with human labels before trusting it at scale.
- Judge-failure semantics — distinguish three outcomes that naive harnesses collapse: scored zero (the artifact is bad), failed closed (the judge errored on a criterion where skipping would vacuously pass), and ungraded (a grading outage — which must never enter the mean, because a dead API key scoring 0.0 is indistinguishable from a terrible output in every downstream aggregate).
- Consistency over ceiling — report passk (all k trials succeed) alongside pass@k, with standard errors; a production agent is judged on reliability, and a “keep the best run” aggregation silently erases the variance you needed to see.
- Mine tasks from production with per-row provenance, and cover both directions — where the behavior should occur and where it shouldn't; one-sided evals produce one-sided optimization.
- Oracle solutions pin the grader — a known-good answer per task must score 1.0, ideally enforced in CI, so the grader and the task generator cannot drift apart without a red test.
- Read transcripts — no score is trusted until someone has looked at what the agent actually did.
The frameworks, compared
| Framework | Paradigm | Measures | Execution | Grading | Eval→RL | Tamper stance |
|---|---|---|---|---|---|---|
| Harbor / Terminal-Bench 2.0 (Laude Institute) | A | terminal/CLI agents, 20+ adapted benchmarks | Docker per trial | unit test / exit code / arbitrary verifier code, optional LLM judge | no | verifier uploaded after the agent phase, isolated |
| SWE-bench Verified | A | GitHub issue → patch | Docker repo checkout | held-out unit tests, binary | no (SWE-Gym forks exist) | human-verified test patches |
| Inspect AI (UK AISI) | A | safety / coding / agentic | Docker, optional K8s | code + model-graded scorers, bootstrap CIs native | no | sandboxed pods |
| HAL (Princeton) | A (meta) | 9 benchmarks, one shared scaffold | shared harness | reuses each benchmark's grader, cost-controlled by default | no | controls the scaffold confound |
| tau2-Bench (Sierra) | A + user-sim | dual-control conversation: agent and a simulated user both mutate shared state | simulated environment | world-state / tool-call checks | no | user-simulator coupled to the environment |
| WebArena / OSWorld 2.0 | A | browser / full computer use | real browser / VM | state-diff of the world | no | VM isolation |
| OpenAI Evals + trace grading | B | multi-turn traces, tool-call correctness | API-level, no sandbox | graders and judges over logged records | RFT — grader-as-reward, OpenAI models only | n/a — grades logs, cannot verify effects |
| verifiers + Environments Hub / SkyRL + OpenEnv | A as training env | RL environments: dataset + parser + rubric + rollout as one artifact | Gym-style isolated envs | rubric-as-reward | yes — the public bridge | environment isolation |
Read the table vertically and one gap jumps out: the paradigm-A frameworks are all eval-only. The eval→RL bridge exists only in the environments-hub ecosystem at the bottom, and in OpenAI's closed RFT. Nothing public combines hermetic effect-verified benchmarking with a training path.
A trace that says “I sent the email” is grading the claim. A verifier that fetched the rendered result is grading the world.
What Harbor actually buys you
It is tempting to read Harbor as “just Docker plus a loop,” and to replace it with a few hundred lines of scripting. The container is not the point. Two properties of the runner are:
The tamper boundary. The agent is an untrusted, reward-seeking process running as root in its container. Anything present during its phase, it can read, edit, or delete — including a grader baked into the image, which is exactly what a naive DIY runner does. Harbor uploads the verifier and its reference data into the environment only after the agent's phase ends. A grader that reads a reference catalog the agent could have rewritten isn't measuring task quality; it's measuring grader-hacking. This attack class is what broke that 16% of public benchmark tasks.
The trial machinery. Per-phase timeouts (an apt mirror stall fails differently from a stuck model), attempts-vs-concurrency semantics, job resume after mid-run death, secret templating so keys never land in committed task dirs, artifact collection, and — critically — a failure taxonomy that keeps errored trials separate from scored-zero trials. Every one of these fails silently when hand-rolled, and silent failures in a measurement system produce confident wrong numbers.
The pattern: one adapter, two projections
The setup this article is really about is a private eval workspace built on both Harbor and Osmosis (an RL post-training platform whose Harbor backend executes Harbor-format tasks inside training rollouts). The load-bearing idea:
A single adapter per eval owns the dataset, the instruction template, and the verifier modules — and projects them into both harnesses. One command generates Harbor task directories for local counterfactual benchmarking; a sibling command generates the platform dataset (one JSON row per task, with the byte-identical instruction text) plus generated copies of the same verifier for the training rollout. The grader that scores the benchmark is provably the same function that emits the RL reward, because both are generated from one source and byte-compare tests fail if a copy drifts.
Around that core, the workspace applies the paradigm-A doctrine with a few implementation choices worth stealing:
- Frozen system-under-test. The real production agent code runs in the trial, but every live dependency is replaced behind ports: file stores instead of databases, JSON catalogs instead of seeded rows, a private CDN instead of a live media service, frozen snapshots instead of live APIs. Infrastructure weather can never move a score or inject noise into a gradient — and when grading infra genuinely fails, the trial errors ungraded rather than scoring zero.
- Reward panels, not scalars. Each trial writes a flat panel: the headline reward, a validity roll-up of every deterministic check, and one key per judge criterion — with never-run criteria absent rather than zero, so “not graded” and “graded 0” stay distinguishable in every aggregate.
- Generated oracles pin the grader in CI. Each task's oracle solution is generated from the row's own grading contract, and a test asserts it scores 1.0. The public benchmarks do this with hours of human review per task; a red test is cheaper and permanent.
- Failure mining, not just task mining. Production defects — an agent reporting success while silently doing something else, invented discount codes, a dropped image reference — become dataset rows and new deterministic checks. A failure isn't considered fixed until a row exists that scores it zero.
- State seeding over transcript replay for multi-turn. To evaluate turn N, materialize the world turn N−1 produced (the artifact, the viewed-entity context) and play a single user message — rather than replaying another model's sampled output, which would make every score depend on that model's behavior.
What the pattern still lacks
Honesty about the gaps, since they map exactly to what each neighboring framework does best:
- Uncertainty and consistency. Point means over small task sets, and a newest-run-wins aggregation that erases variance. Inspect ships bootstrap CIs natively; the passk literature treats every rate as a sample statistic. This is the cheapest, highest-leverage fix.
- Reward-hacking measurement in the RL loop. Where an LLM judge sits inside the training reward, nothing measures the gap between the visible verifier and a held-out variant — the exact instrument the adversarial-hardening work now recommends. A fine-tune faces more verifier pressure than any leaderboard agent, with fewer countermeasures.
- Judge decoupling. One judge call grading all criteria at once, a judge sharing a model family with the system it grades, and a serving-time selection judge coupled to the benchmark grader — three known biases, all fixable with per-criterion isolated calls and a cross-family judge.
- The cost axis. HAL made cost-controlled comparison the default for good reason; a win that costs three times more should be visible in the panel.
- A continuous trace lane. Paradigm B stays manual — humans (or a scheduled agent) panning production traces for failures. The raw material is already captured; grading it continuously is the industrialized version.
- User simulation. Every eval is single-turn over seeded state. tau2's dual-control design is the shape a conversational-orchestrator eval should take — with the caveat that a user simulator is a second model-under-evaluation on the reward path, and must be pinned and calibrated like a judge.
In the two-paradigm frame, the pattern is a fully industrialized paradigm-A shop whose paradigm-B is artisanal. That is both its category and its roadmap — and since no public framework yet occupies the “hermetic benchmark that is also a training environment” cell, it is probably where the public tooling is headed too.