Your Kill Switch May Not Stop Your Agent
The operator stops a procurement agent. Its chat freezes, the worker disappears, and the dashboard t 2026-10-7 07:0:2 Author: hackernoon.com(查看原文) 阅读量:4 收藏

The operator stops a procurement agent. Its chat freezes, the worker disappears, and the dashboard turns gray. Twenty seconds later, a supplier confirms a purchase order.

That sequence is possible without a rogue model, a jailbreak or a broken stop button. The supplier may have accepted the order before the operator acted. A delayed retry may belong to a different worker. A child task may hold its own credentials. Stopping the component you can see does not establish what happened beyond it.

A useful AI agent kill switch needs a precise promise: which new actions it prevents, where that prevention takes effect, and how the team accounts for work already dispatched. The following design uses a hypothetical procurement assistant to build that promise into a revocation timeline and an action inventory. The timings and operation identifiers are illustrative; they describe no production incident.

Give the stop button a bounded promise

Start with the action the operator wants to prevent. For the procurement assistant, the initial scope might be creating supplier orders for one tenant. Reading existing order status remains available to a separate recovery process. Another incident might require stopping every outbound connection from a compromised worker. Those are different operating decisions and need different controls.

Define three observable milestones. Stop requested means an authorized operator has asked for containment. Dispatch blocked means the relevant executor has installed a rule preventing new dispatch authorizations in that scope. Outstanding work accounted for means every known operation crossing that boundary has evidence and an assigned disposition. A user interface should show the milestones separately.

An acknowledgement from the dashboard proves only that the dashboard accepted the request. The executor must acknowledge its own enforcement point. A supplier may have no cancellation feature at all. In that case, the stop contract must explicitly leave accepted orders in the outstanding inventory until their outcomes are established.

The distinction belongs in the product language. Display the affected tenant, tools and runs, followed by remaining operations and unresolved boundaries. If one executor cannot be reached, show partial containment. A single reassuring stopped label would hide the exact gap an incident operator needs to see.

Superorange’s 24-test production readiness review for AI agents on HackerNoon includes cancellation, containment and incident evidence. A stop contract makes those checks concrete for a particular deployment: each promise names a component that can enforce it and a record that can demonstrate it.

Follow the order beyond the agent process

Consider three logical operations belonging to the same run. Order A has reached the supplier. Order B is waiting in the application's dispatch queue. Order C has a pending retry after an earlier response was lost. At the stop request, they require different decisions.

Relative time

Observable event

Consequence for the stop

Before stop

Supplier accepts A; acknowledgement is delayed

A may complete independently of the agent

Stop requested

Operator revokes the run's purchasing authority

Enforcement acknowledgements are still required

Dispatch boundary closes

Executor rejects new authorizations from the revoked run

B stays quarantined; C cannot gain a new dispatch authorization

After boundary closes

Supplier acknowledgement for A arrives

Record the outcome without resuming the run

During reconciliation

Supplier status for C remains unavailable

Preserve uncertainty and assign an owner

The timeline deliberately avoids claiming that the button press and every downstream boundary share one instant. A scheduler, executor and supplier have separate histories. Record their acknowledgements and operation references. Wall-clock timestamps help investigation, but ordering must also come from evidence at the component making the decision.

The Amazon Builders' Library states in Timeouts, retries, and backoff with jitter:

A timeout or failure doesn't necessarily mean that side effects haven't happened.

That observation explains why C cannot be classified as absent merely because its worker timed out. The stop procedure should preserve its original operation identity and consult the supplier's available records. If those records cannot resolve the outcome, the incident still has an unresolved operation. Replacing uncertainty with a fresh order would create another opportunity for an unwanted purchase.

Put revocation where actions are admitted

A prompt telling the agent to stop is useful feedback, but enforcement needs to sit on the path that authorizes tools. In this design, a trusted executor accepts an operation only after checking its authenticated run, tenant, destination, allowed action and current authority. The model cannot set or restore those values through tool arguments.

Give each authority scope a generation number: a version of its permission state. An order admitted under generation 12 carries that generation in its durable record. Revocation advances the authoritative state to generation 13 and closes purchasing. An old worker presenting generation 12 is rejected when it next asks the executor to authorize a dispatch.

The important detail is the race between that check and the send. Reading generation 12, waiting, then sending after generation 13 has been installed is still possible if the check is detached from dispatch admission. Adding a version field to a request does not make that interval disappear.

At an executor you control, serialize the revocation transition with the decision that durably admits a dispatch. An operation admitted before the transition belongs to the outstanding inventory, even if its network send happens later. A decision after the transition must reject the old authority. The receipt should call this an admission boundary; it cannot promise that no previously admitted bytes will leave the machine afterward.

A stronger promise requires a stronger mechanism. You might coordinate a sender barrier and account for every previously admitted send, or use a destination that enforces the same authority version before applying the effect. A remote service that ignores the version provides no such guarantee. Document the actual boundary instead of naming the field a fencing token and assuming the destination honors it.

The admission record must survive a worker crash. Persist the logical operation, permitted payload reference and authority decision before handing work to the sender. If the worker crashes after sending but before storing a response, that durable record keeps the operation visible for reconciliation. If the record cannot be persisted, refuse admission. This ordering does not create a transaction with the supplier; it creates a recoverable account of the local decision. The remaining remote uncertainty is still a separate problem.

Inventory every way work can return

The visible agent loop may be only one producer of actions. A dispatch queue, a scheduled retry, a webhook handler and a child workflow can each bring work back after the original process exits. A stop implementation needs an inventory of these paths before it can claim coverage.

For each path, record who owns it, which identity it executes as, where it checks authority, and whether it can call a supplier directly. Include maintenance scripts and manual replay tools. A carefully guarded executor offers little protection if a retry worker uses the same supplier credentials through a separate connection.

Child operations should retain their parent's run identity and revocation scope unless an explicit policy grants independent authority. Creating a child must not silently turn a revoked purchasing request into a new unrestricted run. If a child genuinely has a separate business purpose, the platform must establish and record that purpose independently.

Queued work needs the same discipline. Quarantine items associated with the revoked scope and require a fresh authorization decision before any later release. Removing an item from one queue does not establish that a leased copy, delayed delivery or retry record no longer exists. Keep enough identity information to recognize those copies when they return.

Late callbacks deserve their own route. An order confirmation is evidence about an existing operation. Store it even while purchasing is stopped, but prevent its arrival from automatically scheduling the next purchase. Separate recording an outcome from continuing the workflow. Otherwise the callback becomes an accidental restart command.

Track delegated capabilities as well as workers. An external job identifier, reusable browser session or previously issued access token may carry authority beyond the originating process. Record where each capability is consumed and how its validity is checked. Revoking the parent account may or may not invalidate a particular delegation; establish that behavior from the service contract and a controlled check. Where there is no recall mechanism, constrain what can be delegated in the first place and include the outstanding capability in the incident scope.

Keep cancellation, revocation and deduplication separate

Cancellation asks ongoing work to stop. Revocation withdraws authority to perform specified actions. Deduplication prevents repeated attempts from applying the same logical effect more than the contract permits. A design needs all three where appropriate, with separate evidence for each.

The gRPC cancellation guide describes a cooperative boundary: the library generally cannot interrupt an application-provided server handler, so the handler must coordinate with cancellation and stop its processing. Propagation also depends on the language and application. A canceled client call therefore cannot stand in for a verified supplier outcome.

An idempotency key solves another problem. If a supplier implements the relevant contract, repeated attempts for one operation can return the existing result without repeating the effect. That does not mean the supplier will reject the first valid attempt after your local stop.

An operation can be perfectly deduplicated and still occur too late for your business requirement.

Amazon's discussion of idempotent APIs examines explicit client request identifiers and late-arriving requests. Check the actual destination's contract: identifier scope, retention, payload mismatch behavior and lookup support. A local identifier alone cannot impose those properties on an external endpoint.

For C, a retry keeps the original logical operation identifier. It also needs current authority to dispatch, regardless of its deduplication properties. While the run is revoked, the proposed policy blocks that new dispatch and permits authorized status lookup. If an earlier attempt had already crossed the admission boundary, account for that attempt separately rather than pretending revocation recalled it.

Make the stop receipt useful to the next operator

The incident record should let another engineer reconstruct the boundary without searching a chat transcript. Begin with the incident identifier, requesting operator, reason, affected scopes and authority generation. Attach each executor's acknowledgement and the revision or sequence it actually installed. Preserve unreachable or stale executors as explicit gaps.

Then list logical operations, with attempts nested beneath them. Each entry needs the business target, run and parent identifiers, admission evidence, destination request reference, latest known outcome, permitted next action and responsible owner. Link to protected evidence rather than copying sensitive payloads into a general incident channel.

For A, the inventory initially says accepted by supplier, final outcome pending. Its late acknowledgement can move it to confirmed created, with a separate decision about whether to request cancellation. For B, a durable rejection or quarantine record establishes that this executor did not authorize a new send after closure. For C, missing external evidence leaves an unresolved outcome.

Do not collapse those entries into a count of failed jobs. Failed jobs measure local execution. The operator needs to know which purchases exist, which may exist, and which pending attempts are prevented from creating another one. An inventory can be complete about known dispatches while still containing unresolved external outcomes.

The receipt also needs a coverage statement. Identify the registries, queues and executor logs used to enumerate work, their observation boundaries, and any missing partitions or retention gaps. If an unregistered worker can dispatch orders, the inventory cannot claim completeness.

Give that gap an owner and keep the containment decision open.

Measure the boundary you can actually enforce

A stop latency measurement is meaningful only with a start event and an end event. Measure request acceptance to each executor's durable enforcement acknowledgement. Separately record the last known pre-revocation admission and any prohibited admission detected afterward. External completion time answers a different question and belongs beside those measurements.

Permission changes may also pass through caches, leases or external identity systems. Inspect their actual behavior instead of assuming a universal revocation delay. A cached permission can remain usable until its implementation requires refresh; a short-lived credential can reduce exposure while leaving already accepted operations untouched. State which mechanism limits each interval.

Choose the response to an unavailable authority service in advance. For purchasing, this proposal refuses new write authorizations when current authority cannot be established. If availability requirements lead to a bounded lease instead, the lease duration becomes part of the containment promise. It is an explicit period of continuing authority, with consequences the owner must accept.

Keep maximum outstanding exposure separate from response speed. A system that admits many parallel purchases can accumulate significant unfinished work before even a fast stop takes effect. Limits on concurrency, operation size and destinations constrain that inventory. Their values should follow the application's tolerable consequences, rather than a generic agent benchmark.

One useful acceptance question is whether the incident owner can enumerate the largest permitted outstanding set. If the architecture allows unbounded fan-out or undiscoverable child tasks, a faster dashboard cannot repair that missing boundary. Fix admission limits and registration before treating stop latency as sufficient evidence.

Also define how overlapping scopes combine. A tenant-level restart must not override a supplier-wide stop imposed by another incident. In this proposal, dispatch requires every applicable scope to permit the action; one active denial is sufficient to block it. The operator interface should identify the denying scope and its owner. Otherwise an engineer may repeatedly restart the tenant, or broaden credentials, while misunderstanding a restriction that is working correctly. Test one overlapping case when verifying the scope logic rather than assuming independent switches compose safely.

Emergency containment sometimes requires terminating workers or cutting their network access immediately. That can be the correct response to suspected compromise. The recovery path should already live outside the affected agent's authority, with independent access to durable evidence and narrowly scoped destination lookups.

Read-only recovery still needs controls. Its credentials should permit only the necessary status queries, its results should go to the incident record, and it should have its own rate limits and access checks. Letting the stopped agent invent recovery calls would restore the very decision-making authority that containment removed.

If a supplier offers cancellation, treat it as a new operation with a separate authorization and receipt. A cancellation request may arrive after fulfillment or may itself have an uncertain outcome. Until the supplier confirms the relevant state, the purchase stays in the inventory. Preserve both the original operation and its cancellation history.

Compensation requires a business decision too. Returning a delivered order, releasing a reservation and canceling an unfulfilled request have different effects. A compensating action may reduce damage without restoring the original state. Assign that decision to an owner who understands the obligation, rather than automatically issuing an inverse tool call.

For teams missing this operating design, Pharos Production's AI consulting governance scope includes rollback procedures and incident response. Those deliverables are relevant when a stop button has been implemented before its recovery responsibilities have been specified. The service scope establishes an available type of work; it does not prove that any particular client's kill switch has passed these checks.

Treat restart as a new authorization decision

Restarting the process should not reopen purchasing. Keep the authority state closed until an authorized owner approves a new scope and generation. A healthy worker proves that computation is available; it says nothing about the disposition of A, B or C.

Require a restart packet that names the incident, current inventory, repaired cause, permitted destinations and remaining exceptions. If only one supplier integration has been repaired, reopen that integration alone. A run-wide green status would conceal the parts that remain restricted.

Old queue entries retain their revoked generation. Do not relabel an entire backlog with the new generation as a convenience. Revalidate each logical operation proposed for continuation against current business intent, target state and unresolved predecessors. B might now be unnecessary because someone placed the order manually during the incident.

C needs particular care. If its outcome remains unresolved, restarting unrelated work does not grant permission to replace it. Keep the relevant business operation blocked or under a separately approved recovery decision. The incident can permit limited service restoration while retaining an unresolved obligation; the record must show both facts.

Separate stopping authority from restarting authority in the on-call policy. An operator may need broad permission to contain suspected harm quickly, while reopening purchases requires the accountable service owner. Name delegates and escalation routes before an incident. A role that exists only during office hours is an operational dependency that the runbook should expose.

Rehearse one stop across the difficult transitions

A useful first drill can be small. Use a test supplier or a controlled simulator, harmless operations and a recorded configuration. Arrange the three operations from the timeline: one accepted with a delayed reply, one waiting for admission and one with an ambiguous earlier result. The purpose is to inspect transitions, not to produce a large model leaderboard.

Trigger revocation while those states exist. Check that the executor records its boundary, queued work cannot obtain new authorization, and late acknowledgements enter the inventory without resuming the workflow. Distinguish an attempted call from a committed business effect in the test evidence.

Then exercise the check-to-send interval deliberately. Pause a test worker after it obtains an admission decision, revoke the scope and release that worker. The expected result must match the declared contract: previously admitted work appears in the outstanding set, or a stronger sender barrier demonstrably prevents it. A test that expects both unrestricted pre-admission and instantaneous recall is testing an impossible promise.

Finally, restart a worker with its old generation and deliver an old callback. Neither event should restore purchasing authority. Reopen a narrow scope through the approved path and verify that quarantined items still require review. Capture the relevant records so another engineer can reproduce the decision without rerunning the entire demonstration.

Close the incident against remaining obligations

A credible completion report can say that purchasing is blocked at all named executors, known dispatches are accounted for, and two supplier outcomes still need investigation. That is useful progress with a defined limit. Hiding the unresolved outcomes behind a stopped process would make the report less accurate.

The kill switch's job is to withdraw specified authority at enforceable boundaries. The incident team's job continues through the work that crossed those boundaries earlier. Design the button, the inventory and the restart decision together, and an operator can explain both what the stop prevented and what still needs attention.

Before calling the feature complete, ask an engineer outside the implementation team to read one stop receipt. They should be able to identify the exact purchasing scope, the enforcement evidence, the operations that might still complete and the person authorized to decide each next step. If those answers require access to the original developer's memory, the receipt is unfinished. A short, reconstructable record makes the control usable when its author is unavailable.


文章来源: https://hackernoon.com/your-kill-switch-may-not-stop-your-agent?source=rss
如有侵权请联系:admin#unsafe.sh