Introduction
Healthcare is deploying AI faster than it builds the controls to run it.
The models improve every quarter. The governance around them does not keep pace. Ask any CTO selling into a health system what slows the deal down. The answer is rarely model accuracy. It’s the security review, the access questions and the “who signs off if this goes wrong” conversation.
For engineering leaders at healthtech startups, this is a hiring story.
The teams who win the next two years will build AI which hospitals trust enough to keep switched on. Those teams need engineers who know how to govern an agent, monitor a model and design an alert clinicians won’t ignore. Very few engineers have done this work in production.
This article covers:
- What new survey data shows about agentic AI in healthcare
- What happens when governance fails, using one of the best-documented cases in health AI
- What separates the deployments which work
- The five governance skills to hire for, and the interview questions to test them
Confidence Is Running Ahead of Control
A new survey from Imprivata puts numbers on the gap. Vanson Bourne conducted the survey on Imprivata’s behalf, reaching 250 U.S. healthcare leaders responsible for identity security or AI strategy. fiercehealthcare
Agentic AI is already live:
- 28% of leaders reported agentic AI already in production, and another 44% said they are piloting AI agents. healthcareitnews
- Only 17% believe existing identity and access management frameworks are sufficient without adaptation. hitconsultant
- 37% said their organisation takes an ad-hoc or unapproved approach to AI agent provisioning, meaning the process of assigning each agent a unique digital identity and permission limits. techtarget
The most telling finding is the contradiction. 86% of respondents said they were fairly confident they could fully control and govern AI-agent actions, yet 72% reported AI tools or agents running without formal IT approval at least occasionally. techinformed
Read those two numbers together. Leaders believe they’re in control. Their own answers say otherwise.
One caveat matters here. Imprivata sells identity and access management products to healthcare organisations and markets a product for securing and governing AI-agent access. Treat the framing with care. The adoption and shadow AI figures still line up with what independent analysts report. KLAS Research’s group director of cybersecurity said KLAS was hearing from healthcare CISOs about agentic AI creating a new identity problem. techinformedtechinformed
Why does this matter to a startup CTO? Because your agent is the one entering their environment. AI agents interact with EHRs, identity systems, clinical apps, medical devices and other technology supporting patient care. Every one of those connections needs an identity, scoped permissions and a monitoring trail. fiercehealthcare
Health systems will start asking for proof of all three before they sign. Someone on your team has to build it.
What Ungoverned AI Looks Like in a Hospital
The best-documented warning in health AI comes from Michigan.
The Epic Sepsis Model runs inside Epic’s EHR and scores patients for sepsis risk. Researchers at the University of Michigan tested it. The retrospective cohort study covered 27,697 adult patients admitted to Michigan Medicine, with 38,455 hospitalisations between December 2018 and October 2019. nih
The results:
- The model did not identify 1,709 patients with sepsis (67%), despite generating alerts for 18% of all hospitalised patients. nih
- At its alert threshold, the model identified only 7% of sepsis patients who were missed by a clinician. healthcareitnews
- The University of Michigan paused the sepsis alerts entirely in April 2020 after complaints about over-alerting. fiercehealthcare
Too many alerts. Too few catches. Clinicians stopped trusting it.
Andrew Wong, the study’s lead author, explained the damage. Alert fatigue “dilutes the importance of other alerts in the system” and contributes to physician burnout. medcitynewsmedcitynews
Here’s the part most people miss. A separate study of the same model reached a different conclusion. Published in Critical Care Medicine, it found the algorithm led to faster antibiotic administration among septic patients and shortened their hospital stays. The lead author of the second study framed the difference well. He argued the validation study asked how the model performs, not how a physician performs with the model. contagionlivecontagionlive
Same model. Opposite outcomes. The difference sat in the deployment.
Two governance decisions shaped the result:
- Thresholds. Hospitals decide which score triggers an alert. Lower thresholds generate more alerts, and higher thresholds risk missing patients. Someone has to own this trade-off with data, not guesswork. medcitynews
- Local validation. The Michigan researchers warned models successful in one setting or time period will not always apply to others. fiercehealthcare
Neither of these is a modelling problem. Both are engineering and governance problems.
What Governed AI Looks Like
Now look at two deployments which worked.
Penda Health, Kenya
Penda Health partnered with OpenAI to study AI Consult, a clinical safety net built into its clinicians’ workflow. The study compared 39,849 patient visits across 15 clinics, with independent physicians rating visits for clinical errors. Clinicians with AI Consult made 16% fewer diagnostic errors and 13% fewer treatment errors. arxiv
The model wasn’t the full story. The authors said the results required a workflow-aligned implementation and active deployment to encourage clinician uptake. Penda backed this with operational discipline. It tracked how often clinicians interacted with AI Consult recommendations and reached out with personalised coaching. arxivopenai
AdventHealth, United States
AdventHealth reported an 80% reduction in time spent on administrative tasks after deploying ChatGPT for Healthcare. The concrete example is utilisation review. Physician advisors spent about 10 minutes per case reading charts and drafting structured rationales. With AI doing the first-pass synthesis, the work drops to roughly two minutes per case, and the clinician still makes the call. openaishellypalmer
Again, the win came from how they ran it. AdventHealth set adoption as a tracked operational KPI and built peer groups by domain rather than running a central training programme. shellypalmer
Rob Purinton, AdventHealth’s Chief AI Officer, named the real challenge: “getting humans to use it safely, consistently, and at scale.” shellypalmer
The pattern across both:
- The AI fits an existing workflow instead of adding a new one
- Someone measures usage and outcomes continuously
- A human keeps accountability for the final decision
- Leaders treat adoption as an operational problem, not a launch event
The model is the commodity. Governance is the differentiator. And governance comes from people.
Independent Evaluation Is Becoming a Discipline
The governance conversation is moving beyond healthcare too.
On 18 September, the AI Evaluator Forum published a public letter calling for independent third-party evaluators at frontier AI companies. More than 100 AI experts signed it, including Geoffrey Hinton. cryptobriefingbnnbloomberg
The letter sets a high bar for independence. The Forum’s existing AEF-1 standard identifies five pillars for credible third-party testing: sufficient access and resources, minimised conflicts of interest, analytic autonomy, transparent methods and results, and protection of sensitive information. ibtimes
Mark Daley, Chief AI Officer at Western University, put it simply. Third-party evaluation is “a principle of engineering safety.” bnnbloomberg
This letter targets frontier AI labs, not health systems. Our read: healthcare will follow the same path. Regulated industries tend to adopt independent testing once the stakes become visible. Expect health systems to ask vendors for evaluation evidence, audit trails and clear answers on who tested the model and how.
Startups who build evaluation infrastructure now will answer those questions in the first security review. Everyone else will scramble.
Five Governance Skills to Hire For
Here’s where this turns into your hiring plan. These skills are scarce, and demand will outpace supply as agentic AI moves from pilot to production.
1. Agent identity and access management
Engineers who treat each agent as its own identity, with scoped, auditable and revocable permissions. Look for people with IAM, zero-trust or healthcare security backgrounds who understand how EHR access works.
2. Production monitoring for model behaviour
Engineers who track drift, alert volumes, override rates and usage patterns after launch. The Michigan case shows why. A model’s behaviour in production matters more than its validation score.
3. Evaluation engineering
People who build evaluation pipelines, test models against local data and red-team failure modes before deployment. This discipline barely existed three years ago. Strong candidates often come from ML infrastructure, QA automation or research engineering.
4. Clinical workflow and alert design
Product engineers who design around clinician attention. They ask how often an alert fires, who receives it and what happens next. Penda’s results came from fitting the tool into the workflow, so this skill drives adoption directly.
5. Audit and compliance infrastructure
Engineers who build logging, traceability and reporting into the system from day one. Under HIPAA, every agent action touching patient data needs a clear record. Retrofitting this later costs far more than building it early.
A few hiring notes:
- Don’t wait for perfect titles. Few candidates hold “AI governance engineer” on their CV. Search for the underlying work: IAM, observability, evals, security engineering in regulated environments.
- Prioritise production experience. Plenty of candidates claim AI fluency. Far fewer have monitored a model in a live environment and fixed something when it broke.
- Hire before the enterprise deal demands it. The first health system security review is the wrong moment to discover nobody owns governance.
- Weigh domain knowledge carefully. Healthcare context helps, but a strong security or infrastructure engineer with fast learning speed often beats a weaker engineer with years of healthtech on the CV.
Interview Questions Worth Stealing
Use these to separate engineers who ship agents from engineers who govern them.
- “Walk me through how you’d set the alert threshold for a clinical prediction model.” Strong answers mention the trade-off between missed cases and alert fatigue, local data and ongoing review. Weak answers stop at accuracy metrics.
- “An agent needs read access to patient records. How do you scope its permissions?” Listen for least privilege, separate agent identities, expiry and audit logging.
- “How would you know your model has started to drift in production?” Good candidates name specific signals: input distribution changes, rising override rates, falling usage.
- “Tell me about a system you shipped which users stopped trusting. What did you change?” This tests ownership. Specific, honest answers beat polished ones.
- “Who signs off before an agent’s permissions expand?” Great candidates design approval into the system instead of relying on goodwill.
One more test for senior hires. Ask how they’d explain a model failure to a hospital’s CISO. The engineers you want will talk about evidence, logs and remediation. The ones to avoid will talk about the model.
Conclusion
Healthcare AI isn’t failing on models. The Penda and AdventHealth results show what well-run AI delivers. The Michigan case shows what happens without the controls around it.
The gap between those outcomes comes down to governance. Governance comes down to the engineers you hire.
The startups who close enterprise deals over the next two years will show health systems three things: who their agents are, what those agents are allowed to do, and how they know the agents are working. Every one of those answers needs a person who has built it before.
Those people are scarce today. They’ll be scarcer next year.
Hiring your first AI governance, evaluation or platform security engineer? These roles are new, and most specs for them miss the mark. Book a role-scoping call with Synaxia Group. We’ll help you define what the hire needs to own in their first 90 days, then find the engineers who’ve done it in production.