An agent is supposed to call lookup_order, then issue_refund, then explain the result. The trace shows exactly that sequence. The evaluation dashboard marks the run as a failure: “required tool missing.”
After an hour of prompt changes, someone inspects the adapter. It renamed issue_refund to refund during normalization, while the contract still expected the original name.
The agent was not failing. The measurement system was. This is a synthetic example, but I encounter the design problem often while maintaining AgentInspect, an open-source local evidence debugger for TypeScript agents. I maintain AgentInspect; it is used here as one concrete implementation of the broader pattern. The argument does not depend on using that library.
Agent evaluations are software. Their parsers, adapters, rubrics, joins, fixtures, and graders can regress. Treating the resulting score as ground truth without testing the evaluator creates a dangerous feedback loop: teams “improve” the agent until it satisfies a broken instrument.
A typical evaluation pipeline is longer than it looks:
agent run
-> framework events
-> trace adapter
-> normalized facts
-> contract or rubric
-> deterministic checks / model grader
-> aggregation
-> dashboard
A defect at any arrow changes the score. Before asking “Why did the agent fail?”, ask four earlier questions:
Consider a required-order check:
type CheckResult =
| { status: "pass"; evidence: string[] }
| { status: "fail"; evidence: string[]; reason: string }
| { status: "not_evaluable"; reason: string }
| { status: "evaluator_error"; errorCode: string };
This is more honest than a boolean. If the trace file is truncated, the adapter does not support the framework event, or a parser throws, false hides the root cause. Worse, some pipelines catch exceptions and return pass to keep CI green.
Fail-open behavior may be appropriate for optional diagnostics, but it must be visible:
function summarize(results: CheckResult[]) {
return {
passed: results.filter(r => r.status === "pass").length,
failed: results.filter(r => r.status === "fail").length,
notEvaluable: results.filter(r => r.status === "not_evaluable").length,
evaluatorErrors: results.filter(r => r.status === "evaluator_error").length,
};
}
A release gate should state what happens for each category. Silence is not a policy.
An agent's regression suite tests the agent. An evaluator's golden suite tests the measuring instrument.
Start with tiny, hand-audited traces:
type GoldenCase = {
name: string;
traceFixture: string;
contractVersion: string;
expected: Record<string, CheckResult["status"]>;
};
const cases: GoldenCase[] = [
{
name: "refund tools occur in order",
traceFixture: "fixtures/refund-happy-path.jsonl",
contractVersion: "refund-v3",
expected: {
requiredOrder: "pass",
forbiddenTools: "pass",
},
},
{
name: "refund before lookup",
traceFixture: "fixtures/refund-reversed.jsonl",
contractVersion: "refund-v3",
expected: {
requiredOrder: "fail",
forbiddenTools: "pass",
},
},
];
Keep fixtures small enough for a reviewer to understand without a dashboard. Include the expected evidence locations, not only the expected label.
Framework events change shape. An SDK upgrade can move a tool name, split streaming events differently, or omit metadata previously available.
Validate normalization at its boundary:
type TraceFacts = {
toolCalls: Array<{
name: string;
callId?: string;
status: "proposed" | "executed" | "failed";
}>;
parseWarnings: string[];
sourceFormat: string;
sourceVersion?: string;
};
expect(normalizeTrace(providerFixture)).toEqual({
toolCalls: [
{ name: "lookup_order", callId: "c1", status: "executed" },
{ name: "issue_refund", callId: "c2", status: "executed" },
],
parseWarnings: [],
sourceFormat: "synthetic-provider",
sourceVersion: "fixture-v2",
});
Do not infer “executed” from “proposed.” Many model APIs emit a requested tool call before your runtime authorizes or runs it. Conflating the two can make unsafe behavior look successful.
Golden happy paths are necessary but easy to satisfy accidentally. Mutate traces and verify that the evaluator notices:
for (const mutation of traceMutations) {
const mutated = mutation.apply(goldenTrace);
const result = evaluate(mutated, contract);
expect(result).toMatchObject(mutation.expectedDetection);
}
If deleting issue_refund does not fail requiredTools, the check is not measuring what its name claims.
Not every evaluation has one golden response. Metamorphic testing checks relationships that should remain true under controlled changes.
Examples:
not_evaluable, not pass.These tests reveal hidden coupling and accidental shortcuts.
Google's Agents CLI evaluation guide separates tool-use quality, trajectory quality, task success, final-response quality, hallucination, grounding, and safety. That taxonomy is useful because one model grader cannot reliably collapse all dimensions into a universal number.
For rubric- or model-based grading:
One grader run is itself stochastic evidence. Repeated grading on a bounded sample reveals variance that a single decimal score hides.
An agent version alone is insufficient. Store:
type EvaluationEnvelope = {
runId: string;
agentVersion: string;
traceSchemaVersion: string;
adapterVersion: string;
contractVersion: string;
evaluatorVersion: string;
graderModel?: string;
graderPromptVersion?: string;
evaluatedAt: string;
};
If a score moves after an adapter release, you can re-evaluate the same raw or sanitized evidence and identify whether the agent changed at all. Never overwrite old scores without retaining their evaluation envelope. Historical dashboards otherwise become a blend of incompatible instruments.
At the time of writing, the current npm release of agent-inspect is 6.19.0 and requires Node.js 20 or newer. Its documented reader API can open a trace and evaluate a contract without re-running the agent:
import {
openTraceFile,
defineTraceContract,
evaluateTraceContractRead,
} from "agent-inspect/reader";
const trace = await openTraceFile("./artifacts/refund-run.jsonl");
const contract = defineTraceContract({
name: "refund-order",
requiredTools: ["lookup_order", "issue_refund"],
requiredOrder: ["lookup_order", "issue_refund"],
forbiddenTools: ["delete_customer"],
});
const result = evaluateTraceContractRead(trace, contract);
console.log(result);
The architectural point is the read-only boundary: captured evidence can be inspected by multiple evaluator versions. That makes evaluator regression testing possible and avoids spending another real-model call merely to test a parser fix. Check exact exports against the version you install; libraries evolve.
Track the health of the measurement path:
trace_parse_error_rate
unsupported_event_rate
not_evaluable_rate
evaluator_exception_rate
grader_timeout_rate
grader_disagreement_rate
evaluation_queue_age
evaluation_version_distribution
These are not agent-quality metrics. They are measurement-quality metrics. Keeping them separate prevents a parser incident from being reported as a model regression.
There is no universal threshold, but the logic should be explicit.
A high-risk workflow might block when:
A low-risk writing assistant may choose a softer policy. Risk, not enthusiasm for metrics, should set the gate.
Before trusting a new evaluation, ask:
The worst evaluator is not one that crashes. A crash gets investigated. The worst evaluator confidently produces the wrong score and sends the team to optimize the wrong system.