FDA Stance on AI Generated Evidence in Drug Applications
FDA outlines seven-step process for validating AI evidence in drug approvals.

For the first time, the FDA's January 2025 draft guidance defines when AI evidence in drug applications is trustworthy. The credibility check has seven steps and focuses on a model's use, not its type, a distinction that's more significant than it initially seems. The agency had authorized more than 1,000 AI-based medical devices before it started formalizing rules for drugs. The FDA is now formalizing what it’s been handling informally for nearly a decade: since 2016, it’s received over 500 drug and biological submissions with AI components, mostly in oncology and neurology.
The build-up wasn't silent, either. In December 2022, Duke Margolis hosted a workshop, the FDA released two papers in May 2023 gathering over 800 comments, and held public workshops in August 2024 and October 2025. Commissioner Robert Califf called the eventual release enabling, not restrictive: "With the appropriate safeguards in place, artificial intelligence has transformative potential to advance clinical research and accelerate medical product development to improve patient care." It reads like a parent giving their kid the car keys along with a set of rules. Meanwhile the AI clinical trial tools market went from $7.73 billion in 2024 to $9.17 billion in 2025, on pace for $21.79 billion by 2030 at almost 19% annual growth. The regulator was writing rules for a market sprinting away from it in real time, and that gap between speed and oversight is the tension the rest of this piece keeps circling back to.
What the January 2025 draft guidance actually covers, and what it deliberately does not
The FDA released the draft guidance "Considerations for the Use of Artificial Intelligence (AI) To Support Regulatory Decision-Making for Drug and Biological Products" (docket FDA-2024-D-4689) on January 7, 2025, accepting comments until April 7, 2025. It’s a draft, so it’s not binding, just suggestions, not requirements. Seven FDA centers and offices joined the document: CDER, CBER, CDRH, CVM, Oncology Center of Excellence, Office of Combination Products, and Office of Inspections and Investigations. With six signatures, the agency intended a single standard, not seven centers developing separate AI rulebooks.
The real tension lies in the scope question. Eligible: AI tools that shape calls on safety, efficacy, or quality. This includes predictive pharmacokinetic modeling that could reduce animal studies, clinical trial design (patient selection, adaptive trials, endpoint analysis), manufacturing quality control, post-marketing pharmacovigilance, and the real-world evidence streams mandated by the 21st Century Cures Act.
Explicitly excluded are AI tools used for early drug discovery, target identification, lead optimization, and internal tasks like drafting submissions or scheduling meetings, as long as they don’t affect patient safety, drug quality, or study result reliability. Hand a chatbot a cover letter to draft and nothing gets triggered. Ask that same model to help set a dosing recommendation, and the full seven-step framework switches on.
Leaving out discovery-phase AI is a mistake, not a justifiable boundary. It leaves decisions that shape everything downstream, like which targets to chase and which patient populations to study, outside any credibility requirement, even though those early calls determine whether the regulated part of the pipeline is even asking the right question. It isn’t just a small boundary decision. It's checking the plane's engine and wings very closely, but not checking who's supplying the fuel. The part about gaps sponsors and critics notice points right to this, since many people raised the same issue during the comment period.
How the seven-step credibility framework is structured
The framework isn't a one-time checklist to complete and forget. It’s a step-by-step process, and the depth a sponsor must reach at each stage hinges on the FDA’s assessment of the model’s risk.
Step one: identify the question being asked. What specific regulatory decision is this model actually informing? If this step’s fuzzy, the rest stays fuzzy too.
Step two is to define the context of use, or COU. It sets the baseline, what goes in, what comes out, and how that affects research or choices.
Step three: measure AI model risk in two ways. How strongly does the model’s output affect the decision (influence), and how serious are the results if it’s wrong (consequence)? The FDA's own example makes this concrete: an AI model could sort patients by adverse-event risk, influencing clinical monitoring decisions. Big sway, big stakes, no shrugging it off.
Step four: create a credibility assessment plan that matches that risk score. The FDA encourages sponsors to engage early in the process.
Run the plan and record all details: data sources, training and testing procedures, and results.
Step six: determine if the evidence is strong enough for this exact context. Not about being broadly reliable, but fit for this particular purpose. A model may predict one thing well, but it's worthless for questions it wasn't tested against, credibility-wise.
The final step: record and submit, and the forms grow with the risk level you picked earlier.
The core idea here is risk-proportionality: high-stakes models like a triage tool get thoroughly vetted, whereas low-risk applications may require just basic evidence. The identical base model, used in two separate contexts, can fall into two totally different regulatory categories. This highlights an issue sponsors often overlook: change management continues beyond approval. Minor updates may follow standard change procedures without additional FDA notice. A moderate change, like retraining on a larger dataset, that affects performance within an already-validated range requires documentation and may need annual reporting. The credibility obligation doesn't stop with approval. It travels with the model for as long as the model stays in use.
The early-engagement architecture the FDA built around the framework
The FDA says it plainly: get in touch sooner. Sponsors who plan to use AI are advised to contact the agency and "set expectations regarding appropriate credibility assessment activities" before investing months in a validation plan that it might reject. Early discussions during the risk assessment and credibility planning stages can help avoid costly corrections later.
The FDA provides multiple engagement pathways for sponsors using AI in regulatory submissions. It looks confusing with all those letters, but the system behind it makes sense. Sponsors pick a program based on the AI model's actual use, so matching the use case to the right option is key, not learning acronyms.
Behind all this sits the CDER AI Council, stood up in 2024, which folded three separate bodies (the AI Steering Committee, the AI Policy Working Group, and the CDER AI Community of Practice) into one voice for AI questions. Bring an untested method to your filing, and odds of a do-over jump. The FDA created these programs to have that discussion sooner, so take advantage of them. A clear process helps sponsors present well-documented submissions. It provides multiple engagement pathways for sponsors using AI in regulatory submissions.
The technical risks the FDA acknowledges the framework must contend with
Any data scientist would immediately recognize the clear concerns the guidance lists. Poor training data leads the list, if the dataset’s too limited, too focused, or secretly lopsided, the model spits out that same flaw every time.
The black-box issue, broken down by Niazi in a 2026 Journal of Chemistry review, gets divided into three parts often lumped together. Explainability means figuring out why a model gave a certain result, mainly for developers and reviewers examining the details. Interpretability checks if a clinician can understand the output using their everyday language. Transparency checks if the design, training data, and validation methods can actually be reviewed. Even if a model shares all its training data, it might still not make clinical sense, and sponsors often make this mistake by just giving a data dictionary for "explainability." Mixing up these three ideas is why credibility plans usually need revisions.
A model's real-world data after approval can quietly diverge from its validation data, posing a unique risk, especially for post-market surveillance. The guidance also notes a quieter danger: AI can spin multiple believable lines of reasoning that all point to the wrong answer, and that very believability makes reviewers drop their usual doubts. Call it the algorithmic equivalent of a con artist who's repeated his story so often he now believes it. The guidance keeps a human involved for this reason, even when models validate properly.
Reproducibility completes the list. Reproducing AI predictions for clinical trials is often challenging, and the large datasets used for training bring additional concerns about data security and privacy. Common recommendations include collecting representative data, testing across demographic groups, involving human oversight, and addressing bias early in development. This isn't novel. It’s the basic care any cautious statistician would use, now written into FDA requirements. Mistake even one part, and the sponsor’s proof falls apart from the ground up.
Where the guidance leaves gaps that sponsors and critics have flagged
Let's start with the discovery-phase exclusion, as it's the gap drawing by far the most opposition. Target identification and lead optimization are completely outside the framework, but those early decisions determine which groups are studied and which questions are asked later, exactly where the framework applies. The European Medicines Agency's framework includes the entire product lifecycle. This mismatch could create different compliance requirements for sponsors running programs in both jurisdictions.
Data governance and privacy considerations are not extensively addressed in the draft. Some analysts have warned that this gap may become more urgent over time. The guidance also doesn't address measurement: it says credibility assessment should match the level of risk, but doesn't specify how to measure risk or credibility to get FDA approval. Sponsors have to set their own limits and hope the agency agrees, which is strange since guidance is meant to stop that kind of guessing.
The January draft barely mentions generative AI, a major oversight considering how quickly large language models have become part of research workflows. How do you properly assess risks for a model with outputs that vary each time it runs? How do you fit a constantly changing ensemble model into a system designed for static ones? What about a vendor's AI service on infrastructure the sponsor doesn't control? There's still no clear answer to any of that, and saying the seven-step framework already does cover it is unrealistic.
The Niazi study and other academics have called for risk-based explainability requirements. Different sponsors may have varying capacities to implement validation requirements. The FDA has asked the public whether the seven-step approach fits sponsors’ real experiences, whether more guidance is needed for lifecycle maintenance and pharmacovigilance, and if numeric thresholds for enough evidence would work in practice. These questions remain open and unresolved.
Where the guidance stands now and what comes next for sponsors
The comment period closed April 7, 2025, and as of early 2026 the FDA has signaled final guidance sometime in the second quarter, folding in what came back during the comment window. On January 14, 2026, the FDA and the European Medicines Agency each released ten shared principles for using artificial intelligence responsibly in drug development, covering everything from early research to post-market stages. For sponsors working across countries, that alignment is more important than it first appears, as it begins to fill the precise discovery-phase gap the U.S. draft didn't cover alone.
But it's not done yet. Academic reviewers also want that alignment to include China's NMPA and Japan's PMDA, and global sponsors must navigate a patchwork until it does. The final guidance is expected to clarify credibility evidence standards at each risk level, provide detailed rules for LLMs and generative AI, and maybe offer separate guidance on lifecycle maintenance and pharmacovigilance, based on FDA comment-period questions.
Sponsors acting now, before the final text arrives, should focus on four key steps. Compare each AI use in your pipeline to current rules now, to avoid rushing later when the final guidance comes out. Talk to the FDA early using the named programs (C3TI, CID, ETP, CATT) before finalizing a credibility assessment plan that could later require changes. Integrate the context-of-use definition and risk assessment directly into model development, not as a later, unwanted documentation task. Start tracking updates on day one, retraining or tweaking the model later means extra paperwork, even after approval.
The underlying message is that this framework is intentionally designed to scale from low-risk applications to models that directly triage patients for life-threatening conditions. Sponsors who include that scaling logic in model development, not just as a compliance add-on after the science is complete, won’t be caught off guard when the final version arrives. Those still gambling on discovery-phase AI remaining exempt forever should revisit the EMA's lifecycle framework before doubling down.
Sources
- A Critical Review of the FDA’s Draft Guidance on Artificial Intelligence in Drug and Biological Product Regulation - Niazi - 2026 - Journal of Chemistry - Wiley Online Library
- Artificial Intelligence for Drug Development
- Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products
- Key takeaways from FDA’s draft guidance on use of AI in drug and biological life cycle | DLA Piper
- fda.gov
- troutman.com
- jonesday.com

