AI Diagnostic Tool Procurement Evaluation Checklist for Hospitals
Hospitals need a framework for evaluating AI tools that evolve after purchase.

Buy an MRI machine and whatever gets installed in March is still running in December. Buy an AI diagnostic tool, though, and by the second Tuesday of the fiscal year you could be looking at functionally different software, because the algorithm underneath can get retrained on new data without anyone wheeling a new box through the loading dock. That one fact breaks most procurement templates before they even get opened. AI diagnostic tools sit across medical device rules, clinical evidence standards, data privacy law, and health IT interoperability requirements at the same time, and nobody has built a template that covers all four at once, so hospitals are mostly stitching one together as they go.
Every vendor demo I've sat through sounds finished. The gap between how finished it sounds and how finished it actually is turns out to be where most of the real risk lives. The FDA tried to deal with the moving-target problem in its December 2024 guidance on Predetermined Change Control Plans, or PCCPs: vendors spell out ahead of time how they intend to modify an algorithm after clearance, so nobody gets blindsided by a silent update pushed out over a weekend. Tidy on paper, but in practice the hospital's job doesn't end at signature, it changes shape, and most procurement offices haven't reorganized around that yet, largely because nobody told them they'd need to.
Clearance and clinical validation measure two different things, no matter how often a sales deck blurs the line between them. A device can be FDA-authorized and still have zero prospective, outcomes-based trials behind it. That's just how the pathway works, and treating "cleared" as shorthand for "proven" skips the one step that would have actually told you something.
Then there's the paperwork nobody thinks about until it's six months overdue. A 2025 HHS proposed rule says organizations using AI tools have to fold those tools into ongoing risk analysis, same as any other system touching patient data. Procurement opens a file, and that file stays open long after the closing date, whether anyone remembers to check it or not. Most published research on hospital AI covers the planning stage, the part where everyone gets to feel smart around a conference table. Far less exists on the doing, the stretch where a tool either holds up in production or falls apart quietly, and hospitals are largely improvising that part right now. A checklist keeps you from improvising blind, which, given what's at stake, is a lower bar than it should be.
Regulatory status: what clearance pathways actually tell you, and what they don't
Ask which pathway a device came through: 510(k), De Novo, or PMA. Each one carries a different evidence bar, and a device cleared through the lowest doesn't automatically compare to something that cleared the highest. The AHA has noted that FDA rules for software as a medical device require premarket testing of safety and efficacy, but "tested" means wildly different things depending on which door the vendor walked through, and vendors are not exactly rushing to explain which door that was.
If the tool is headed anywhere outside the U.S., check for a CE mark or the relevant national authorization, and don't assume it carries over just because the box says Europe on it. The EU AI Act, published in July 2024, requires that training, validation, and test datasets for high-risk AI be relevant, sufficiently representative, error-free, and complete, and it demands documented bias mitigation and ongoing performance monitoring on top of the existing Medical Device Regulation. A tool sold in Europe is clearing two bars at once.
Domestically, Federal health IT rules have continued to add transparency and algorithmic decision-making requirements for AI built into certified health IT. If the tool plugs into the EHR, those requirements apply, and no exceptions get carved out for how small the vendor is, no matter how earnestly their sales rep insists otherwise.
One flag worth circling on any RFP: a vendor who leans hard on "FDA-cleared" language but has no filed PCCP. That doesn't disqualify them outright, but it does mean they're making an implicit promise that the algorithm stays frozen in place, and that promise gets harder to keep the longer the product sits on the market. Ask directly, since vendors rarely volunteer the answer unprompted.
Four things need confirming before this stage closes: clearance pathway, PCCP status, international authorizations where relevant, and a post-market surveillance plan actually on file. Leave one blank, and the decision on whether to move forward has effectively already been made.
Clinical evidence quality: reading validation studies the way a procurement team should

A tool can carry FDA clearance, get called "validated" in every slide the vendor sends over, and still have never run a prospective trial anywhere near a setting that resembles the hospital buying it. This is worth checking every single time, not just when something feels off.
Keep IDx-DR in your back pocket as the reference point. The diabetic retinopathy tool ran a prospective trial of 900 patients and came back with 87.2% sensitivity and 90.7% specificity for detecting referable retinopathy. That's what solid prospective validation looks like: a real number, from a real trial, on a population large enough to mean something. Set that against a review of 151 AI imaging devices approved as of late 2021, where only 64.2% clearly stated their use of clinical data in the FDA summary at all, and the average validation set sat at 799 patient cases. That number is a floor, and plenty of tools land right at it, or somehow below it.
Study designs fall into a rough hierarchy. Prospective, consecutive-patient studies sit at the top; they cut down on selection bias and test the tool against real-world variability instead of a curated sample built to flatter the numbers. Retrospective multicenter studies work fine as supporting evidence, but they shouldn't carry the whole case alone. Single-site studies on a highly curated dataset belong at the bottom, and they should get flagged as insufficient for a hospital-wide rollout no matter how clean the numbers look on the page.
Ask for metrics broken out by patient subgroup, not lumped into one tidy aggregate: sensitivity, specificity, area under the ROC curve, positive and negative predictive value, and performance across age, sex, race and ethnicity, and comorbidity burden. A vendor who can only hand over the aggregate number has revealed a gap in what they've bothered to check, and it's usually the thing you most needed to know.
Post-market evidence matters too, and payers look at outcomes data independent of whatever the clearance process required. In 2023, the FDA granted a New Technology Add-On Payment for an AI sepsis detection tool based on evidence it improved patient outcomes in practice. That's the kind of evidence that should make a procurement team feel good about signing, and its absence should make someone ask why it's missing, out loud, before the meeting ends.
Four fields, tracked honestly: study design tier, sample size measured against the roughly 800-patient benchmark, subgroup reporting present or absent, and post-market findings available or not. That tells you more than any case study the vendor's marketing team assembled for the pitch deck, no matter how nice the slides look.
Algorithmic bias and health equity: the evaluation step most procurement teams still skip
Sit with this number for a second: 74% of hospitals using predictive AI in 2024 checked their tools for bias, according to AHA Center for Health Innovation data. Flip that around, and a real chunk of hospitals didn't check at all. For a tool making or informing diagnostic decisions about actual patients, that gap should bother whoever's signing the contract, and not just the compliance office two floors down.
Why does this bite specifically in diagnostics? Training data that doesn't reflect a hospital's actual patient mix can bake existing healthcare disparities directly into the tool's output. An algorithm trained mostly on one demographic doesn't generalize to another just because the underlying disease looks similar on paper. Disease doesn't present identically across populations, and a model that's never seen the variation won't catch it when it counts most.
So ask the uncomfortable questions before anyone signs anything. What's the demographic makeup of the training and validation data, and does it resemble the hospital's own patients? What bias mitigation is actually documented, not just promised on a call? (This is a formal requirement under the EU AI Act for high-risk tools, for whatever that's worth to a buyer who doesn't operate in Europe.) Will the vendor commit in writing to regular third-party bias audits? Can they show performance broken down by race, ethnicity, sex, age, and some proxy for economic status, instead of one clean number that smooths over everything underneath it?
Standard SaaS contracts weren't written with any of this in mind, so it has to get added by hand, clause by clause, sometimes over the objections of a legal team that's never negotiated one of these before and isn't thrilled to be learning on the job now. Vendor agreements should require validation on data that actually reflects diversity, and they should bind the vendor to an audit schedule that doesn't happen once at signing and then quietly disappear.
Here's a test that costs nothing and takes five minutes: ask the vendor for performance numbers on the specific subpopulations that make up the largest share of your hospital's patients. Silence is the answer if they can't produce that number, and it's not a subtle one.
EHR integration and interoperability: where most AI implementations quietly break down
A tool can post excellent numbers in validation and still fall apart the moment it meets a hospital's actual EHR. Messy data formatting and thin documentation wear down model performance once it's live, and the gap between "worked in the study" and "works on Tuesday at 2pm in the ED" usually comes down to plumbing: the kind that's boring, unglamorous, and nobody budgets for until it's too late.
The plumbing isn't cheap either. Hooking an AI tool up to a legacy system that isn't FHIR-compliant can add anywhere from $20,000 to well over $100,000 in custom APIs, middleware, and data translation work, a figure that can eclipse what the vendor spent building the AI in the first place. That cost belongs in the procurement conversation up front, not something IT discovers six weeks into implementation while asking pointed questions nobody has a good answer for.
Set the technical baseline in the RFP itself: HL7 FHIR compliance at R4 or better, SMART on FHIR app launch capability, confirmed compatibility with the hospital's current EHR version and data model. About 80% of hospitals using AI get it straight from their EHR vendor, per the AHA's IT Supplement, while 52% use a third-party tool, and some run both at once. The EHR-vendor route usually means simpler integration but tighter lock-in; the third-party route buys flexibility but dumps a heavier interoperability burden on the hospital's own team. Neither path is free of tradeoffs, and pretending otherwise is how the surprises show up later, always at the worst possible time.
Federal health IT rules fold AI transparency requirements into certified health IT, so confirm the integration layer actually meets that bar instead of assuming it does just because the EHR itself is certified. Those are separate claims, even though vendors treat them as interchangeable more often than they should.
There are workflow questions too, and they too often get punted to the go-live meeting instead of asked in the RFP. Where does the AI's output actually show up for the clinician: an alert, a dashboard, a recommendation buried three clicks deep in a tab nobody opens? What happens when the output doesn't render cleanly in the EHR interface? Who, by name, owns the integration once the EHR vendor pushes its next platform update? "The vendor will handle it" is a deferral dressed up as an answer, and it should get treated as one.
Data privacy, security, and vendor contract terms that protect the hospital after signing
Healthcare data breaches averaged $9.77 million each in 2024, the highest of any industry for the fourteenth year running. An AI diagnostic tool that touches protected health information widens that attack surface, and about 47% of breaches in 2025 trace back to supply chain attacks. An AI vendor is a node in that supply chain whether anyone labels it that way or not, and the vendor's security posture becomes the hospital's exposure the moment the contract gets signed, not sometime later when it's convenient to think about it.
Some contract terms shouldn't be negotiable. A signed Business Associate Agreement, no exceptions carved out for a "trusted" startup that swears it's different. Encryption at AES-256 minimum, backed by third-party certification, SOC 2 Type II or HITRUST, rather than a vendor's own word that they take security "seriously." That claim means nothing on a contract. There should also be an explicit clause barring the vendor from using patient data to train its general-purpose or commercial models without separate, explicit consent, a term that's oddly absent from most standard SaaS agreements because most standard SaaS agreements were never written with PHI in mind at all.
The regulatory ground is shifting underneath this too. HHS proposed the first major update to the HIPAA Security Rule in twenty years on January 6, 2025, removing the old split between "required" and "addressable" safeguards and tightening encryption and resilience expectations across the board. Any AI vendor contract signed now should anticipate where that rule lands, not just where it currently stands.
Breach notification deadlines matter more than they sound like they should. HIPAA technically allows up to 60 days, but best practice is pushing hospitals toward 24 to 72 hour windows written directly into vendor contracts. Worth pinning down explicitly whether an AI output error that accidentally exposes PHI even counts as a reportable incident under the contract's language, because that's exactly the edge case that gets argued about after the fact.
OCR has been clear on one point: covered entities can't hand off HIPAA compliance to a vendor. If the vendor gets breached, the hospital is still the one making the notification calls, regardless of whose server the data happened to sit on. That fact alone should reframe how much scrutiny the security section of the RFP actually deserves; in practice it usually gets less than the clinical section, which is backwards.
Governance, monitoring, and what the hospital must own after deployment begins
Signing the contract starts the workload; it doesn't end it. AI diagnostic tools need ongoing performance monitoring, drift detection, and periodic revalidation, and the PCCP framework has turned that from a nice idea into something closer to a regulatory expectation with actual teeth.
Before go-live, a hospital needs a governance structure already standing, not sketched out after something goes wrong. That means a named clinical champion and a named technical owner for each tool, a defined review cycle with quarterly audits at minimum, a clear escalation path for when performance slips or a bias signal shows up, and a process for checking every vendor algorithm update against what the hospital's PCCP review actually expects to see.
Model drift matters because of patient safety, not because it's a box on a compliance form. A model retrained on new data without regulatory review is, functionally, a different device wearing the same product name on the login screen, and hospitals need a way to catch that shift before it becomes a safety problem, not after someone notices the numbers look strange in a retrospective chart review six months later. Confirm the vendor has a real post-market surveillance plan, and get specific about what data they'll actually share and how adverse events or unexpected outputs get reported to the FDA. Vague reassurance here isn't worth much.
None of this matters if clinicians won't use the thing anyway. Only 16% of clinicians currently use AI tools to help make clinical decisions, according to Elsevier's 2025 "Clinician of the Future" survey, a strikingly small number given how many hospitals report having these tools integrated already. That gap is central to the whole effort, and governance planning has to include training, feedback loops, and honest tracking of adoption, because a tool that clears every regulatory and clinical bar but sits ignored on a dashboard has accomplished a very expensive nothing.
Running the checklist under real procurement conditions
Timelines slip, even for well-resourced systems that know exactly what they're doing and have done this before. NHS England's 2025 study of AI procurement for chest diagnostics across 66 Trusts found contracts expected by November 2023 weren't actually signed until March through September 2024, and clinical deployment planned for December 2023 didn't start at some sites until May 2024. That gap says something about how much genuine diligence this kind of purchase demands, even when everyone in the room wants it to move fast and the vendor keeps mentioning the quarter-end discount as though that changes the math.
Build the evaluation time into the project plan from day one, then, instead of treating it as the thing that gets squeezed the moment someone upstairs asks why this is taking so long. A workable order: confirm regulatory status and clearance pathway before any vendor demo happens. Review the clinical evidence and bias attestation before anything gets shortlisted. Finalize interoperability requirements and security contract terms before negotiation starts in earnest, and get the governance structure and post-deployment monitoring plan actually built, not just discussed once in a meeting, before go-live.
This diligence work marks the difference between buying a tool and buying a commitment that keeps changing shape long after the ink dries. Slowing down is how you catch the version of this that gets signed under pressure, in a market moving faster than most hospitals' own governance can keep pace with. That's a slowdown worth choosing on purpose, not one you back into after the fact.


