Architecting Resilient AI Systems for Healthcare
Clinical software does not get the slack most other software gets. When a recommendation is wrong and nobody catches it, the consequence is not a bug ticket to be picked up in a later sprint, it is a decision about someone's care that was made without the right information in front of the person making it. That reality shapes how we build, and it pushes the design work away from optimising for how a model performs on average and towards understanding how it behaves when it is wrong.
The first commitment is that architectural decisions get written down at the time they are made rather than reconstructed from memory months later. Which model version is running in production, what confidence threshold sends a case to a human reviewer, and what the system does when an input does not fit the expected pattern all belong in a record that exists before anyone asks for it. When a regulator or a clinical team wants to know why the system behaved a particular way, the answer needs to be already on hand.
Testing carries more weight in this work than it does in most. An update can raise overall accuracy while quietly getting worse at catching something rare, and an aggregate accuracy score will not surface that. Every update therefore runs against a fixed set of difficult and atypical cases before it goes anywhere near a patient, and the results on those cases matter more to us than the headline number.
Failure also has to be visible to the person using the system. A model that is not confident should say so and either flag the case or hand it back to a health worker, rather than returning a clean-looking answer that happens to be wrong. An answer presented without any signal of uncertainty gives the reader no reason to question it, which means it usually will not be questioned.
The final call stays with the person rather than the model. A health worker can override any recommendation, and those overrides are tracked, because a clinician disagreeing with the system tells us something real about where its judgement and theirs diverge. That divergence is worth watching over time, both as a safety measure and as a signal about where the model needs work.
The work does not end at launch either. Patient populations shift, seasons change what is common, data entry habits drift, and a model that performed well on day one can degrade months later without a single line of code changing. Monitoring for that has to be designed into the build rather than added once something has already gone wrong.
None of these practices are unusual on their own. Most serious engineering teams keep decision records, test properly and monitor after deployment. What changes in healthcare is the size of the gap that opens when any one of them is skipped, because the cost is not measured in defects but in decisions made about someone's health without the safeguard those decisions needed.