A few years back, a hospital in the Midwest rolled out a sepsis prediction model across its ICU. Within weeks, nurses were quietly muting the alerts. Not because the model was wrong exactly, but because it fired so often that people stopped trusting it, and eventually stopped listening to it altogether. That's the story that keeps coming up whenever someone tells me AI is "ready" to make clinical calls on its own. It usually isn't a story about bad math. It's a story about bad design around the math.
I think that distinction matters more than most people give it credit for. Building a model that predicts patient deterioration, flags a suspicious mass on a scan, or recommends a drug dosage is, honestly, the easier half of the problem now. There's plenty of tooling for that. The harder half, the part that barely gets discussed at conferences, is figuring out where the human belongs in that loop and how much authority to actually hand over.
Healthcare has this reputation for being slow to adopt technology, and engineers coming from other industries sometimes read that as backwardness. I'd push back on that a little. Medicine is one of the few domains where a false negative can mean someone dies and a false positive can mean unnecessary treatment, unnecessary cost, or a patient losing trust in the system entirely. Neither error is "acceptable loss" the way it might be in, say, an ad-recommendation engine.
There's also the distribution shift problem, which gets underestimated constantly. A model trained on data from one hospital system, with its particular patient demographics, lab equipment, and charting habits, doesn't necessarily generalize to a hospital three states over. IBM's Watson for Oncology ran into exactly this wall years ago: recommendations trained largely on synthetic cases and one institution's treatment patterns didn't translate cleanly elsewhere, and clinicians noticed pretty quickly. That episode is worth remembering, not to bash the effort (plenty of smart people worked on it), but because it's a clean example of what happens when the system assumes its own outputs are portable truth rather than one input among several.
So no, full automation isn't really on the table for most consequential clinical decisions, at least not yet, and maybe not for a long while. What is on the table, and what's actually being built right now, is human-in-the-loop design done properly. Which brings me to what "properly" even means here, because it's thrown around loosely enough to have lost some meaning.
A lot of early clinical AI deployments technically had a human in the loop. A doctor would glance at the model's output and click "approve." That's not really human-in-the-loop, that's human-as-liability-shield. The clinician has thirty seconds, an alert fatigue problem of their own, and zero visibility into why the model said what it said. Under those conditions people default to trusting the machine, which defeats the entire point of having them there.
Real human-in-the-loop design means the system is architected so a person can meaningfully evaluate, question, or override the model's output, with enough context to actually do that job. That has implications for the whole pipeline, not just a checkbox at the end.
Here's roughly how that pipeline tends to look in the systems I've seen work reasonably well:

Patient data comes in from the EHR, monitoring devices, labs, whatever the relevant sources are. The inference engine produces a prediction, but critically it's paired with a confidence or uncertainty estimate rather than a bare output. That confidence score feeds an escalation router, which is really the piece doing the interesting work architecturally. Low-uncertainty, low-stakes predictions might just get logged quietly. High-uncertainty predictions, or predictions in categories that are inherently high-stakes regardless of confidence (dosing, sepsis, code status changes), get routed to an actual clinician review interface, one that shows the reasoning, not just the verdict. The clinician's decision, and any override, gets logged. That log doesn't just sit there for compliance either, it feeds back into monitoring and eventually retraining.
Notice what this buys you. The clinician isn't reviewing every single output, which would recreate the alert fatigue problem from that sepsis rollout I mentioned earlier. But they're also not being handed a black box only when it's convenient for the vendor's liability position. The routing logic is doing triage on when human judgment is actually worth the interruption cost.
I'd argue the escalation router deserves way more design attention than it usually gets. Threshold tuning here isn't a one-time calibration exercise, it's closer to an ongoing negotiation between sensitivity and clinician trust. Set the threshold too conservatively and you're back to alert fatigue. Set it too loosely and you risk missing the case that actually mattered, which is obviously the worse failure mode in a hospital.

What I like about laying it out as a flow like this is that it forces you to be explicit about categories that get treated as always-escalate regardless of model confidence. A model might be 98% confident about a medication interaction and still, by policy, require a pharmacist's eyes on it if the interaction falls into a certain severity class. That's a business rule sitting on top of the statistical one, and pretending the statistics alone should decide is, I think, a category error people make when they're more comfortable with the model than with the domain.
The override path matters just as much as the escalation path, maybe more. When a clinician disagrees with the model, that disagreement needs to be captured with a reason, not just a binary "overridden" flag. Those reasons are gold for catching model drift early, and honestly for catching bad models before they cause real harm. If clinicians are consistently overriding a particular type of recommendation, that's a signal worth investigating before it becomes a lawsuit or, worse, a patient safety incident.
I want to be a little careful here because "explainable AI" gets oversold constantly, and I've contributed to that literature myself, so I'm aware of the gap between the marketing and the reality. SHAP values and attention maps give clinicians something to look at, sure, but they're not the same as a causal explanation, and presenting them as such risks a false sense of understanding that might be worse than no explanation at all.
What actually seems to help, from what I've observed, is showing the clinician which input features drove the prediction in relatively plain terms (this patient's lactate trend, this specific EKG segment), alongside similar historical cases and their outcomes, rather than trying to fully "explain" a neural network's internal reasoning. It's a more modest goal, and it's achievable. Overpromising interpretability tends to backfire once clinicians realize the explanation doesn't hold up under scrutiny.
The FDA's framework for AI/ML-based Software as a Medical Device, and the EU AI Act's classification of most clinical AI as high-risk, both push toward exactly the architecture described above: logging, human oversight, and post-market monitoring aren't bureaucratic add-ons, they're the mechanism by which a system earns the right to be trusted incrementally rather than all at once. Engineers building in this space who treat regulatory requirements as a checkbox to satisfy after the fact tend to end up rebuilding half their system later. It's cheaper, in my experience, to design the audit trail and the escalation logic in from day one.
If you're designing something in this space, a few things worth internalizing early: build the confidence estimator as a first-class component, not an afterthought bolted onto model outputs. Design your escalation thresholds with actual clinicians in the room, not as a purely statistical exercise, because alert fatigue is a real clinical outcome with real consequences. Log overrides with reasons, and actually review that log periodically rather than letting it accumulate as dead weight. And resist the temptation to oversell explainability features as more definitive than they are.
None of this is glamorous work. It won't get you a demo that wows a room the way a flashy diagnostic accuracy number does. But it's the difference between a system clinicians quietly route around and one they actually rely on. The human in the loop isn't a limitation on the technology, or a temporary stage before "full AI" arrives. In medicine specifically, it might just be the whole point.