How to Regression-Test the Systems Evaluating Your AI Agents
An agent is supposed to call lookup_order, then issue_refund, then explain the result. The trace sho 2026-10-5 16:4:18 Author: hackernoon.com(查看原文) 阅读量:1 收藏

An agent is supposed to call lookup_order, then issue_refund, then explain the result. The trace shows exactly that sequence. The evaluation dashboard marks the run as a failure: “required tool missing.”

After an hour of prompt changes, someone inspects the adapter. It renamed issue_refund to refund during normalization, while the contract still expected the original name.

The agent was not failing. The measurement system was. This is a synthetic example, but I encounter the design problem often while maintaining AgentInspect, an open-source local evidence debugger for TypeScript agents. I maintain AgentInspect; it is used here as one concrete implementation of the broader pattern. The argument does not depend on using that library.

Agent evaluations are software. Their parsers, adapters, rubrics, joins, fixtures, and graders can regress. Treating the resulting score as ground truth without testing the evaluator creates a dangerous feedback loop: teams “improve” the agent until it satisfies a broken instrument.

Every Score Has a Provenance Chain

A typical evaluation pipeline is longer than it looks:

agent run
  -> framework events
  -> trace adapter
  -> normalized facts
  -> contract or rubric
  -> deterministic checks / model grader
  -> aggregation
  -> dashboard

A defect at any arrow changes the score. Before asking “Why did the agent fail?”, ask four earlier questions:

  1. Did we capture the relevant event?
  2. Did we normalize it correctly?
  3. Did the evaluator apply the intended rule?
  4. Did the report preserve uncertainty and errors?

Do Not Collapse Missing Evidence Into Failure—or Success

Consider a required-order check:

type CheckResult =
  | { status: "pass"; evidence: string[] }
  | { status: "fail"; evidence: string[]; reason: string }
  | { status: "not_evaluable"; reason: string }
  | { status: "evaluator_error"; errorCode: string };

This is more honest than a boolean. If the trace file is truncated, the adapter does not support the framework event, or a parser throws, false hides the root cause. Worse, some pipelines catch exceptions and return pass to keep CI green.

Fail-open behavior may be appropriate for optional diagnostics, but it must be visible:

function summarize(results: CheckResult[]) {
  return {
    passed: results.filter(r => r.status === "pass").length,
    failed: results.filter(r => r.status === "fail").length,
    notEvaluable: results.filter(r => r.status === "not_evaluable").length,
    evaluatorErrors: results.filter(r => r.status === "evaluator_error").length,
  };
}

A release gate should state what happens for each category. Silence is not a policy.

Create a Golden Suite for the Evaluator

An agent's regression suite tests the agent. An evaluator's golden suite tests the measuring instrument.

Start with tiny, hand-audited traces:

type GoldenCase = {
  name: string;
  traceFixture: string;
  contractVersion: string;
  expected: Record<string, CheckResult["status"]>;
};

const cases: GoldenCase[] = [
  {
    name: "refund tools occur in order",
    traceFixture: "fixtures/refund-happy-path.jsonl",
    contractVersion: "refund-v3",
    expected: {
      requiredOrder: "pass",
      forbiddenTools: "pass",
    },
  },
  {
    name: "refund before lookup",
    traceFixture: "fixtures/refund-reversed.jsonl",
    contractVersion: "refund-v3",
    expected: {
      requiredOrder: "fail",
      forbiddenTools: "pass",
    },
  },
];

Keep fixtures small enough for a reviewer to understand without a dashboard. Include the expected evidence locations, not only the expected label.

Test the Parser and Adapter Separately

Framework events change shape. An SDK upgrade can move a tool name, split streaming events differently, or omit metadata previously available.

Validate normalization at its boundary:

type TraceFacts = {
  toolCalls: Array<{
    name: string;
    callId?: string;
    status: "proposed" | "executed" | "failed";
  }>;
  parseWarnings: string[];
  sourceFormat: string;
  sourceVersion?: string;
};

expect(normalizeTrace(providerFixture)).toEqual({
  toolCalls: [
    { name: "lookup_order", callId: "c1", status: "executed" },
    { name: "issue_refund", callId: "c2", status: "executed" },
  ],
  parseWarnings: [],
  sourceFormat: "synthetic-provider",
  sourceVersion: "fixture-v2",
});

Do not infer “executed” from “proposed.” Many model APIs emit a requested tool call before your runtime authorizes or runs it. Conflating the two can make unsafe behavior look successful.

Mutation Testing Exposes Weak Checks

Golden happy paths are necessary but easy to satisfy accidentally. Mutate traces and verify that the evaluator notices:

  • delete a required event;
  • reverse two tools;
  • duplicate a side effect;
  • change an argument after approval;
  • add an unexpected tool;
  • truncate the final line;
  • corrupt a timestamp;
  • remove a parent or handoff reference; and
  • rename an event field as a simulated SDK change.
for (const mutation of traceMutations) {
  const mutated = mutation.apply(goldenTrace);
  const result = evaluate(mutated, contract);

  expect(result).toMatchObject(mutation.expectedDetection);
}

If deleting issue_refund does not fail requiredTools, the check is not measuring what its name claims.

Not every evaluation has one golden response. Metamorphic testing checks relationships that should remain true under controlled changes.

Examples:

  • changing whitespace must not change extracted tool order;
  • reordering unrelated metadata must not change a contract result;
  • adding a benign explanatory sentence must not authorize a tool;
  • increasing a payment amount must not preserve an old approval grant;
  • replacing a user name with another valid name must not alter a tool allowlist; and
  • removing all evidence must produce not_evaluable, not pass.

These tests reveal hidden coupling and accidental shortcuts.

Model Graders Need Calibration, Too

Google's Agents CLI evaluation guide separates tool-use quality, trajectory quality, task success, final-response quality, hallucination, grounding, and safety. That taxonomy is useful because one model grader cannot reliably collapse all dimensions into a universal number.

For rubric- or model-based grading:

  • maintain a human-labeled calibration set;
  • record grader model and configuration;
  • rerun a stable subset after prompt or model changes;
  • measure agreement and disagreement, not only average score;
  • blind the grader to irrelevant experiment labels;
  • keep deterministic facts outside the model grader; and
  • route uncertain or high-impact disagreements to review.

One grader run is itself stochastic evidence. Repeated grading on a bounded sample reveals variance that a single decimal score hides.

Version the Whole Evaluation Envelope

An agent version alone is insufficient. Store:

type EvaluationEnvelope = {
  runId: string;
  agentVersion: string;
  traceSchemaVersion: string;
  adapterVersion: string;
  contractVersion: string;
  evaluatorVersion: string;
  graderModel?: string;
  graderPromptVersion?: string;
  evaluatedAt: string;
};

If a score moves after an adapter release, you can re-evaluate the same raw or sanitized evidence and identify whether the agent changed at all. Never overwrite old scores without retaining their evaluation envelope. Historical dashboards otherwise become a blend of incompatible instruments.

A Concrete Read-Only Evaluation Boundary

At the time of writing, the current npm release of agent-inspect is 6.19.0 and requires Node.js 20 or newer. Its documented reader API can open a trace and evaluate a contract without re-running the agent:

import {
  openTraceFile,
  defineTraceContract,
  evaluateTraceContractRead,
} from "agent-inspect/reader";

const trace = await openTraceFile("./artifacts/refund-run.jsonl");

const contract = defineTraceContract({
  name: "refund-order",
  requiredTools: ["lookup_order", "issue_refund"],
  requiredOrder: ["lookup_order", "issue_refund"],
  forbiddenTools: ["delete_customer"],
});

const result = evaluateTraceContractRead(trace, contract);
console.log(result);

The architectural point is the read-only boundary: captured evidence can be inspected by multiple evaluator versions. That makes evaluator regression testing possible and avoids spending another real-model call merely to test a parser fix. Check exact exports against the version you install; libraries evolve.

Observe the Evaluator as a Dependency

Track the health of the measurement path:

trace_parse_error_rate
unsupported_event_rate
not_evaluable_rate
evaluator_exception_rate
grader_timeout_rate
grader_disagreement_rate
evaluation_queue_age
evaluation_version_distribution

These are not agent-quality metrics. They are measurement-quality metrics. Keeping them separate prevents a parser incident from being reported as a model regression.

What Should Block a Release?

There is no universal threshold, but the logic should be explicit.

A high-risk workflow might block when:

  • any deterministic safety contract fails;
  • evaluator errors exceed a tiny reviewed tolerance;
  • required evidence is missing for a material portion of the cohort;
  • grader calibration falls below an agreed range;
  • an adapter version is unrecognized; or
  • results cannot be reproduced from retained evidence.

A low-risk writing assistant may choose a softer policy. Risk, not enthusiasm for metrics, should set the gate.

The Instrument Deserves Its Own Test Plan

Before trusting a new evaluation, ask:

  1. Which exact facts does it consume?
  2. How do we know those facts were parsed correctly?
  3. What golden cases prove pass, fail, and not-evaluable behavior?
  4. Which mutations must the evaluator catch?
  5. How stable is any model grader?
  6. Which versions produced the score?
  7. Can we rerun evaluation without rerunning the agent?
  8. How does missing evidence affect release decisions?

The worst evaluator is not one that crashes. A crash gets investigated. The worst evaluator confidently produces the wrong score and sends the team to optimize the wrong system.

References


文章来源: https://hackernoon.com/how-to-regression-test-the-systems-evaluating-your-ai-agents?source=rss
如有侵权请联系:admin#unsafe.sh