Your Agent Retries Because It Can't Tell a Timeout From a Failure
TL;DR: My agent hit a timeout, treated it as a failure, and re-ran a task that had already finished. 2026-10-4 04:48:42 Author: hackernoon.com(查看原文) 阅读量:2 收藏

TL;DR: My agent hit a timeout, treated it as a failure, and re-ran a task that had already finished. It correctly diagnosed every error it saw and kept retrying anyway. Understanding a failure does not constrain what an LLM does next, so the constraint has to live in the execution layer, and that layer needs a state most retry logic lacks: unknown.

I have since turned the fix into a small open-source gate, with a demo you can run in a few seconds: Harness-Engineering on GitHub. This article explains why it works the way it does.

The task had already completed

I am building an agent that reviews automated test coverage. It executes test scripts, then compares what they exercise with the application design, the source code, and the acceptance criteria of the user stories. Its job is to find requirements that look implemented but have no meaningful test, and tests that pass without checking the behavior they claim to cover.

For a while it behaved well. Then an execution was interrupted, and the agent retried. The retry failed, so it retried again. Across iterations it correctly named what was going wrong: timeouts, locked records, a missing VPN connection. None of that changed what it did next. It kept trying to execute the task.

When I finally stepped in, the task had completed long before. The first timeout had hidden a successful run. The agent's retry then hit a locked record, left behind by that first run, and it kept retrying, creating new records each time. It ran four full times before I noticed. It was retrying into a mess of its own making, and no retry wrapper in my code was involved: the loop was the model's own reasoning.

That is why I built a kill switch. The more useful lesson came afterward: stopping an agent is necessary, but an agent that executes work needs a way to tell a failed action from an action whose result is unknown.

A diagnosis is not a constraint

This is the part that surprised me. The agent was not confused about the errors. It could tell me that the VPN was down and that a record was locked. Its explanation was accurate and its behavior did not follow from it.

An LLM's reasoning about a failure is text. Nothing forces the next action to be consistent with that text. If the loop says "retry until done" and the only tool on offer is run_tests, the agent will keep calling run_tests, with an articulate commentary on why it keeps failing.

So the rule I now work by: anything that must hold regardless of what the model concludes has to be enforced outside the model. Prompts can ask for caution. Only the execution layer can refuse.

And there is a second problem underneath. Even a perfectly obedient agent cannot decide correctly if the system hands it only two outcomes.

A timeout is an observation, not an outcome

From the agent's side, two very different situations look identical:

  • The operation never completed.
  • The operation completed, but its response never arrived.

Both produce a timeout. Only the first suggests that another execution might be needed. In the second, a retry repeats work that already happened.

This matters because executing a test is not a read-only operation. Depending on the suite, it can create records, change application state, acquire locks, trigger background processing, or call another service. "Run the test again" can mean "perform every side effect again."

The fix is to stop collapsing the result into success or failure. Every attempt gets a third possible state, unknown, and unknown has its own exits: you cannot go from unknown straight back to submitting. You have to reconcile first.

Sakti Bagchi's image-44d948

The only way from an unknown attempt back to submitting runs through a lookup by the original run key, and through a person when the lookup cannot settle it.

In the repo these are the ledger states submitting, complete, failed, unknown, needs_inspection and released. Only failed, within a retry budget, and released may start again.

Write the intent down before you act

My agent's assignment sounds like one task, "review test coverage". Operationally it is five different actions, and each needs its own retry policy:

Action

Safe to repeat?

What makes it safe

Read code, tests, designs, stories

Yes

Record which revisions were read

Start a test run

No

A stable run key, deduplicated by the runner

Poll status, fetch results

Yes

Read-only tools, separate from the start tool

Test setup and cleanup

Only with protection

Per-run fixtures, repeatable operations, or a verified reset

Save findings

Yes

An upsert keyed by review ID

The dangerous row is the second one. A retry is only safe if the same intended run keeps the same identity, so the key must be derived from what the run is, not generated fresh on each attempt:

from harness_attempt import key_for, execute, reconcile

# The key comes from what the run is, never from when it ran.
key = key_for(review_id, suite, environment, application_commit, test_commit)

# The intent is written to the ledger before run_suite is called.
# A TimeoutError inside it is recorded as unknown, not as failed.
execute(root, key, payload, run_suite, actor="coverage-agent")

# The same key is now refused until the outcome is settled.
reconcile(root, key, probe, actor="coverage-agent")

The order matters. The intent is written to the ledger before the call, because the runner may accept the submission and the response may be lost before you ever see a run ID. The stored run key is the only handle that lets recovery find that run later. A checkpoint kept in the agent's conversation history will not survive an interruption, so it lives in a store the agent cannot rewrite.

Starting a run has to be atomic. A separate "look up, then start" sequence is not enough, because two callers can both see that no run exists, so the ledger takes an exclusive lock before it decides. If the service you call can also deduplicate by key, pass the same key through and the gate becomes your second line of defense. If neither side can look a run up, an uncertain submission stops for a person: that is the needs_inspection state.

After a timeout while waiting for results, the next permitted action is a status check, never a new start. That is why start_test_run, get_test_run_status and get_test_run_results are separate tools: a timeout changes how you observe a run, it does not authorize a new one.

Two smaller points. Stable identity does not make everything inside a test safe to repeat: setup may create a record before failing, and cleanup may time out after deleting one. For tests that mutate state I require at least one of per-run fixtures, repeatable operations, or a verified reset, otherwise the agent refuses another attempt until a person checks the environment. And reads should record the revisions they used, so a resumed review never quietly combines yesterday's test run with today's source code:

evidence = {
    "application_commit": application_commit,
    "test_commit": test_commit,
    "design_revision": design_revision,
    "story_revision": story_revision,
}

Put the gate where the actions happen

A retry budget limits damage, but it cannot decide whether repeating an action is valid. Three attempts at an unsafe operation are still unsafe. What I needed was a gate in the executor that closes on specific conditions, whatever the model proposes:

def execution_permission(context):
    if context.kill_switch_enabled:
        return Deny("Execution stopped by operator")
    if context.previous_attempt.outcome == "UNKNOWN":
        return Deny("Reconcile the previous attempt first")
    if not context.vpn_connected:
        return Deny("Required connectivity is unavailable")
    if context.record_lock_unresolved:
        return Deny("Inspect the existing lock before continuing")
    if not context.environment_ready:
        return Deny("Verify test environment state")
    return Allow()

def execute(action, context):
    permission = execution_permission(context)
    if permission.denied:
        return blocked(permission.reason)
    return dispatch(action)

Each condition closes the gate for a different reason. A missing VPN is a prerequisite failure: repeating a dependent action will not restore connectivity. A locked record is evidence that something else may still be running, or that an earlier attempt left state behind, so the right response is to inspect ownership, not to try again. And an unknown outcome is its own state that must never be folded into a generic "failed" bucket.

The sketch above is simplified. The version in the repo is scripts/harness_attempt.py, and it refuses with ten named codes, among them ATTEMPT_OUTCOME_UNKNOWN, ATTEMPT_KILL_SWITCH, ATTEMPT_PRECONDITION_FAILED and ATTEMPT_BUDGET_EXHAUSTED. Each code has a defect fixture that must trip it and a clean fixture that must not, so the gate itself is tested, not just used.

These are the same three errors my agent diagnosed correctly and then ignored. The difference now is that the diagnosis no longer matters: the executor refuses regardless of what the model concludes.

The kill switch belongs in the same place. An agent saying "I have stopped" is useful feedback; the execution service refusing further actions is the enforceable control. Stopping future actions and cancelling an active run are also separate controls. Cancellation needs support from the runner, plus verification of the state it leaves behind.

In the repo, stop writes a stop file that refuses all new work, and only an operator can lift it. Work that already finished can still be replayed while it is on, because returning a recorded result changes nothing.

When the system cannot determine what happened, it should preserve the evidence and stop mutating anything. The request to a person should be specific enough to act on:

Test submission timed out. Execution outcome could not be confirmed.
Further execution is blocked.

Inspect the existing run and test environment, then mark the attempt
completed, still running, or safe to execute again.

That gives the operator a concrete decision, and it stops the agent from reading one more failed status check as permission to start over.

The agent sees the same decision in a form it cannot misread. This is the refusal from the repo's wrapper (the key is shortened):

UNKNOWN key=3b8c...: no result after 600.0s. The work may have finished. Do NOT run it again.
REFUSED ATTEMPT_OUTCOME_UNKNOWN: the earlier attempt of 3b8c... timed out and may have finished
  next: Do not run it again. Find out what happened: reconcile --key 3b8c... --probe "COMMAND"

The operator's side is a separate command that refuses to run as the same actor who started the attempt, so the agent cannot settle its own doubt.

Unverified is not the same as a gap

This matters beyond preventing duplicate work. If the agent cannot establish whether a test completed, it cannot use that run to judge coverage. A disrupted environment, an unresolved lock, or a repeated mutation can contaminate the evidence behind every conclusion that follows.

So the report has to keep the distinction instead of guessing:

finding = {
    "criterion": acceptance_criterion,
    "coverage_status": "UNVERIFIED",
    "reason": "Execution outcome remains unknown",
}

An unverified criterion is different from a confirmed coverage gap. Reporting the first as the second sends a team off to write tests they may already have.

One honest limit: the gate does not label findings for you. It only makes sure an unknown outcome is visible. Your reporting code still has to apply the rule.

What changed

I reproduced the incident in a simulation: a lost response, a lock left behind by the first run, and a scripted agent that reads each error correctly and runs the task again. It runs once without the gate and once through it. The counts below are exact for that simulation.

No gate

Gate

Gate, nobody can look

Suite executions

4

1

1

Records created

4

1

1

Failed attempts

3

0

0

Task confirmed finished

no

yes

yes

In the last column nothing can see the effect, so the gate cannot settle the question itself. It stops and a person decides, which is the intended behavior.

Three things this does not show. The agent is scripted, so it says nothing about how a live model behaves. It counts executions and records, not minutes or tokens, so I cannot yet put a number on time or tokens saved, only say that every execution it prevents is one nobody waits for and no model has to read the output of. And it covers one failure shape. Running a real agent through the gate and publishing the result is the next step.

Run it yourself

The demo needs Python 3.10 or newer and nothing else installed:

git clone https://github.com/sakti1977/Harness-Engineering
cd Harness-Engineering
python3 -m examples.gate.timeout_demo

It ends with TIMEOUT DEMO PASSED. docs/attempts.md covers the command line, what the gate checks, and its limits, including that it only controls work that goes through it.

Before you ship an agent that executes work

Write down which actions are safe to run twice, which can be reconciled through a durable identifier, and which must stop for inspection. Then enforce those decisions where the actions happen, not in the prompt.

Otherwise, a lost response can become permission to do the same work again.

Sources


文章来源: https://hackernoon.com/your-agent-retries-because-it-cant-tell-a-timeout-from-a-failure?source=rss
如有侵权请联系:admin#unsafe.sh