Most production systems already know when something is wrong.
Logs. Alerts. Database rows. Monitoring feeds. Runbooks. APIs that can restart a service, change config, retry a workflow, or fix bad state.
Those pieces usually sit in different rooms.
So a break still turns into a human scavenger hunt. An alert fires. Someone opens a dashboard, greps logs, checks a table, hunts docs, maybe compares it to an incident from six months ago. After enough of that, somebody names a cause and hits an API.
I want an agent that walks that loop itself: investigate, decide whether it understands enough to touch production, run a controlled fix, then check that the system actually came back.
That is a self-healing agent. Not another chatbot that guesses for an on-call engineer.
I have spent my time in distributed systems, cloud-native backends, retrieval-augmented generation, LLM orchestration, autonomous agents, and tool-based AI. The design below is those pieces wired into one operational path.
This will sound backwards in an agent article. I still would not let an LLM decide when remediation starts.
That job is for deterministic systems.
Say a transaction gets stuck in an intermediate state. You may already have a monitor like:
IF transaction.status = "PROCESSING"
AND transaction.age > threshold
AND downstream_service = healthy
THEN create remediation_event
Or an observability rule that notices a weird combo:
error_rate > baseline
AND database_latency = normal
AND dependency_health = healthy
The exact condition is not the point. Rule-based triage is the front door.
Rules are good at known operational states. Predictable. Cheap. Auditable. Testable.
LLMs earn their keep later, when the situation is less of a boolean.
I treat the system as stacked layers:
Detection Layer
Rule-Based Triage
Agentic Investigation
Controlled Remediation
Verification
The rule says something odd happened. The agent figures out what it means.
A typical incident bot pulls a few log lines and dumps them on an LLM. That is almost never enough.
A real incident is spread across different kinds of evidence: relational or transactional data, application logs, system events, monitoring feeds, old incident records, architecture docs, runbooks, API schemas, config history.
That is why hybrid RAG matters here.
Do not treat RAG as "search some docs and paste them into a prompt." Operational RAG has to pull from sources that are not the same type of evidence. A SQL row is not a runbook paragraph. A log event is not a metric.
Suppose triage reports:
Observation:
Return workflow has remained in PENDING state longer than expected.
The agent should not guess yet. It collects.
Inspect transactional state first:
SELECT status,
retry_count,
last_updated,
workflow_stage
FROM workflow_records
WHERE workflow_id = :id;
Maybe the workflow reached stage four and never recorded completion. That is the what. It is still not the why.
Another agent pulls logs around that timestamp. It might see:
Validation completed
Routing request submitted
Downstream response timeout
Retry scheduled
Retry handler exited
Now you have a possible failure path.
A monitoring agent looks at service health. Maybe the downstream dependency is healthy now.
That matters. If the dependency is still down, a retry is wasted, or worse.
A documentation agent searches ops guidance for that signature.
If validation is complete and the downstream service has recovered, resume the workflow from the routing stage rather than restarting the full transaction.
Now you have observations plus a procedure. That is a lot more to work with than asking an LLM "Why is this workflow stuck?"
You could dump all of this into one huge window and hope one model sorts it. I would not.
Give agents narrow jobs:
The Data Agent answers questions about current application state. The Log Agent reconstructs what happened. The Documentation Agent finds known remediations. The Health Agent checks whether dependents are safe to touch. The Diagnosis Agent merges the evidence. The Remediation Agent picks an approved action. The Verification Agent checks whether that action actually fixed it.
Easier to follow. Easier to debug. If the diagnosis is wrong, I can look at what each agent retrieved instead of unwinding one giant chat.
The interesting part is leaving "I think I know" and actually fixing it.
An LLM should not get open production access. Expose approved capabilities through Model Context Protocol, or MCP.
MCP is the agent's controlled door to the outside. Not arbitrary network. Specific tools, for example:
get_workflow_state()
get_service_health()
retrieve_logs()
retry_workflow()
resume_workflow()
rollback_change()
verify_transaction()
The agent can think widely. It can only act through those functions.
Keep that split.
The model says:
I believe the safest remediation is resume_workflow(id=1234)
The MCP tool layer decides whether that function exists, whether the arguments are valid, whether this agent may call it, and whether extra policy checks apply.
That is not an assistant with opinions. That is an operator with a short tool list.
"Zero-touch" gets used like a slogan. I mean something smaller.
For specific, well-understood incident categories, the system can finish the loop without waiting on a person.
For example:
resume_workflow(id=1234) through MCPverify_transaction() and confirms the workflow left PENDINGIf health is still bad, or the runbook does not match, or verification fails, it stops and pages a human. Zero-touch is for the cases you already know how to fix. It is not a blank check.
The loop is the product: detect with rules, investigate with hybrid RAG, act only through MCP, prove the system recovered. That is the self-healing agent I would actually ship.
I would not wire the system so every diagnosis fires an action.
The agent should score whether the evidence is strong enough to justify a fix.
Picture four signals lining up:
Database state: workflow stalled after validation
Logs: retry handler failed after dependency timeout
Service health: dependency currently healthy
Runbook: resume workflow when dependency recovers
Those agree. Autonomous execution can make sense.
Now flip it:
Database state: inconsistent
Logs: incomplete
Service health: unknown
Runbook: two possible remediation paths
That is an escalation, not zero-touch.
A working autonomy policy:
|
Confidence |
Action |
|---|---|
|
High confidence |
execute automatically |
|
Medium confidence |
propose action for human approval |
|
Low confidence |
gather additional evidence or escalate |
You keep automation where the story is clean. You stop short when it is not.
A common automation mistake: treating an API response as proof the problem is gone.
Say the remediation agent calls:
resume_workflow(1234)
and gets:
HTTP 200 OK
The API accepted the request. That is all. The system may still be sick.
The verification agent has to look at what happened next. For example:
workflow.status == COMPLETE
AND error_rate returned to baseline
AND no duplicate transaction created
AND downstream acknowledgement received
Close the incident only then.
The architecture is a loop, not a one-way pipeline:
OBSERVE
↓
TRIAGE
↓
INVESTIGATE
↓
DIAGNOSE
↓
ACT
↓
VERIFY
↓
HEALTHY?
├── YES → close
└── NO → re-investigate
That last check is the difference between self-healing and a bot that just pokes APIs.
Give a system more autonomy and you need a better paper trail.
Every automated remediation should record:
Keep the retrieved context on the incident too. An engineer reviewing later should see what the agent saw, not a summary of a summary.
That matters in large enterprises, where ops work can bump into compliance and governance.
Autonomy with no observability is another outage waiting to happen.
Do not start with the nastiest incidents. Self-healing should earn more freedom over time.
Good first cases usually have:
A stuck workflow is a decent start. A corrupted distributed data model across five services with fuzzy ownership is not.
Skip "let AI fix everything." Look for the operational problems where people already run the same investigation and the same fix, then automate that path with the brakes on.
Once that holds up, widen the scope.
A lot of enterprise AI still helps people find documents. Fine. Useful, even.
The more interesting work starts when retrieval, reasoning, tools, and verification sit on the same path.
RAG gives context. Multiple agents split the investigation. MCP gives controlled access to real application capabilities. Rules start the process on something you can audit. Verification closes it.
Wire those together and an event can move from:
Alert
↓
Engineer
↓
Investigation
↓
Fix
to:
Alert
↓
Agentic Investigation
↓
Verified Remediation
People still write the policies, tools, fences, and allowed remediations. The AI does the repetitive operational reasoning inside those fences.
That is the self-healing story I actually buy. Not software magically sewing itself back together. Software that can finally connect what it already knows about its own state with the procedures and APIs it needs to recover.
Once a system can observe, understand, act, and verify with some reliability, an alert stops being a ping. It is the start of the repair.