I have spent eighteen years keeping enterprise systems running, and most of the last several of them on a global business intelligence platform that sits under a Fortune 500 company's reporting, forecasting, and planning. Every line of business touches it. When it is late, people notice within minutes.
The platform is not one thing. It is a relational warehouse, an in-memory column store, a cloud data platform, a streaming layer, and a distributed SQL engine, plus the schedulers and integration middleware holding them together. Each of those components ships with good monitoring. Each of them will tell you, in detail, that it is fine.
And they can all be fine while the number on the executive dashboard is wrong.
That gap is the actual operational problem in large analytics platforms, and I do not think most teams have named it yet. We bought monitoring per engine. We never bought observability of the answer.
Here is the failure I have seen more often than any outage.
A source system changes a field. Not dramatically. A code that used to be six characters becomes seven, or a currency column starts arriving with different rounding, or an upstream team adds a row type nobody told us about. Ingestion succeeds. The transformation runs. The load completes. Every job goes green.
Two days later a regional finance lead says their forecast looks off by a few percent, and now four teams start reading logs backwards from a spreadsheet.
Nothing broke. That is the point. The pipeline did exactly what it was built to do, which was move data, not verify meaning. Our alerting was instrumented on the verb, not the noun. We alerted on "did the job run" when the business cares about "is the number right and did it land before the 6 a.m. review."
Every time I have gone back through one of these incidents, the signal was there. It was just scattered. A row-count anomaly in one system's logs, a latency spike in the streaming layer, a retry pattern in the scheduler, a job that finished successfully but took forty minutes longer than its usual eleven. Four systems, four log formats, four different on-call engineers, none of whom had the other three windows open.
The traditional answer to this is a runbook and a senior engineer. It works, and I have relied on it for years, but it has a hard ceiling.
The correlation work is not intellectually hard. It is wide. You are asking somebody at 2 a.m. to hold six mental models at once, remember what normal looks like for each of them, and notice that a mild anomaly in one place plus a mild anomaly in another place equals a real problem. Humans are bad at that specific task, and they get worse the more systems you add. Meanwhile the number of systems only goes up.
Adding people does not fix it either, because each new person has to build their own map of a platform that changes weekly. I have watched capable engineers spend their first six months just learning where things live.
This is where generative AI earned its place in my operations work, and it is not the place most people expect.
Most enterprise conversations about AI right now are about producing things. Write the code, write the query, write the summary.
The higher-value use in an always-on platform is reading. Specifically, reading heterogeneous machine output across systems that were never designed to be read together, and telling a human where to look.
We started pointing LLM-based workflows at exactly that. Pull the scheduler state, the engine logs, the streaming lag metrics, and the recent job history for a given business process. Ask what is unusual relative to that process's own history. Produce a short, ranked, plain-language hypothesis with the evidence attached.
A few things I did not expect:
Format-agnostic reading is the whole trick. The reason cross-system correlation was expensive was never intelligence. It was a translation. Five systems, five vocabularies for the same underlying event. A model that can read all five dialects without a parser per source removes the actual cost.
Baselines beat thresholds. Our thresholds were guesses that ossified. "Alert if runtime exceeds sixty minutes" was set by somebody in 2019 for a data volume that no longer exists. Comparing a run against the shape of its own recent history catches the slow drift that static thresholds are structurally incapable of catching.
The output has to be a hypothesis, not a verdict. The first version we built stated conclusions. Engineers correctly stopped trusting it the first time it was confidently wrong. The version that stuck says "these three signals are unusual together, most likely explanation is X, here are the logs." That framing survived contact with skeptical senior engineers because it respects them. It hands over evidence and lets them judge.
The measurable result was not headcount. It was time-to-locate. The expensive part of an incident was never the fix. It was the forty minutes of five people asking each other whether it was their layer.
I run this playbook on application security too, and the logic is identical.
Vulnerability scanning is not a detection problem anymore. Scanners produce plenty of findings. The problem is that a large application portfolio generates far more findings than any team can act on, and the ranking that comes out of the box is generic. It does not know which of your applications is internet-facing, which one handles regulated data, which one is scheduled for decommission in two quarters.
Severity without context is just a queue. And a queue nobody can finish is a queue nobody trusts.
What actually changes the posture is correlating findings with what the system does, what it touches, and what it is worth. That is the same cross-source reasoning problem as the pipeline case, wearing different clothes.
Some hard-won opinions, stated plainly.
Instrument the business outcome, not the job. If your alerting cannot tell you that the number is wrong while every job is green, you are monitoring your infrastructure and hoping about your data. Define your SLA on the answer's correctness and arrival time. Everything else is a leading indicator.
Do not let the AI act. Ours reads, correlates, and recommends. A human decides. I have no interest in an autonomous agent restarting jobs on a platform where a bad restart corrupts a downstream forecast that somebody presents on Thursday. That line may move eventually. It has not moved for me yet.
Log quality is now a feature. The moment an LLM became a consumer of our logs, sloppy logging turned into an operational cost we could feel. Consistent identifiers across systems went from a nice-to-have to the thing that determines whether correlation works.
Start with your worst recurring incident. Not a platform strategy. One incident type that keeps coming back and burns senior time every month. Build the correlation for that one, measure whether it shortens time-to-locate, then decide if it generalizes. Enterprise AI programs die of scope long before they die of technology.
I do not know where the ceiling is on this.
The current setup handles known-shaped problems well, meaning things that have gone wrong before in some recognizable form. The novel failure, the one where a vendor changes something upstream that nobody documented, still comes down to an experienced engineer with a hunch. I have not seen a model develop a hunch. I have seen models produce very plausible sentences about the wrong subsystem.
So I am not going to claim this replaces operational judgment. What it does is stop wasting that judgment on tasks that are wide but shallow, which is most of what an on-call engineer actually spends their night doing.
Eighteen years of running always-on systems has taught me one thing that keeps proving itself: the cost of an incident is dominated by uncertainty, not by repair. Anything that shortens the uncertainty is worth more than it looks on a slide.
Your engines are probably fine. The question is whether anyone is watching the seams.