Regulation explainerLife Sciences
How to validate AI in GxP systems under Part 11, Annex 11 and draft Annex 22
GxP AI validation means showing, with evidence an inspector can follow, that a computerized system containing a model does what its intended use requires. The core rules are 21 CFR Part 11 and EU GMP Annex 11, with the draft Annex 22 adding model-specific expectations for critical manufacturing uses. This explainer turns them into a classification step, a validation sequence and an evidence file to keep for each model.
On this page
- What this explainer covers, and where to take legal advice
- Why machine learning strains classic computer system validation
- The rule set for AI in GMP manufacturing and quality systems
- Classify each AI use before planning a single test
- A validation sequence for one model, from intended use to release
- The evidence file to keep for every validated model
- When supplier, cloud and model-vendor evidence can be relied on
- A visual inspection model on a fill-finish line, from risk assessment to release
- Questions and answers
- Sources
What this explainer covers, and where to take legal advice
Why machine learning strains classic computer system validation
Classic computer system validation assumes logic someone wrote down: requirements are specified, code implements them, and scripted tests show that a given input produces the specified output. A machine learning model weakens each link. Its behavior comes from training data, its output may be a probability, and its accuracy can decay when inputs drift away from what it learned.
Explaining one decision is harder too, because an image classifier does not say why it rejected a vial. The draft EU GMP Annex 22 responds by narrowing what may be used in critical applications at all: models that do not adapt during use and that return identical outputs for identical inputs4.
The rule set for AI in GMP manufacturing and quality systems
Six instruments shape validation today. Two are drafts, so check their status before citing them in a validation plan.
Title 21 CFR Part 11: electronic records and signatures
United States (FDA)Applies whenRecords required by FDA predicate rules, such as batch records, are created, modified, kept or transmitted electronically1.
- Validate systems for accuracy, reliability and consistent intended performance.
- Keep secure, computer-generated, time-stamped audit trails and control access.
EU GMP Annex 11: Computerised Systems (EudraLex Volume 4)
European UnionApplies whenA computerized system supports GMP activities; the January 2011 revision is the text published in EudraLex2, and a revised draft went to consultation from July to October 20253.
- Apply risk management across the lifecycle and assess suppliers.
- Validate, keep audit trails, evaluate periodically and control changes and configuration.
Draft EU GMP Annex 22: Artificial Intelligence
European Union (consultation draft, not yet adopted)Applies whenAI models are used in critical applications with direct impact on patient safety, product quality or data integrity in manufacturing medicinal products or active substances4.
- Only static, deterministic models in critical applications; generative AI and large language models are excluded from them.
- Intended use and acceptance criteria approved by a process expert before testing, at least matching the process replaced.
- Representative, stratified test data kept independent of development, plus feature attribution, confidence logging, change control and monitoring.
FDA guidance: Computer Software Assurance for Production and Quality System Software
United States (FDA, final guidance)Applies whenSoftware supports medical device production or the quality system; FDA announced the final guidance in September 20255. Drug GMP is outside its scope, but its reasoning transfers.
- Base assurance on each feature's intended use and the harm its failure could cause.
- Use unscripted testing and lighter records where risk is low.
ISPE GAMP 5: A Risk-Based Approach to Compliant GxP Computerized Systems (Second Edition)
Industry guidance (international)Applies whenUsed as the reference lifecycle for GxP computerized systems; the July 2022 edition added appendices on AI and machine learning and discusses Computer Software Assurance6.
- Apply critical thinking to decide which testing adds assurance, reusing supplier evidence where adequate.
FDA draft guidance: Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products
United States (FDA, draft guidance)Applies whenA model produces information supporting decisions about a drug's safety, effectiveness or quality, manufacturing included; discovery and purely operational uses are excluded7.
- Define the question of interest and context of use, rate model risk from influence and consequence, then plan and document a credibility assessment.
Classify each AI use before planning a single test
Classify the specific task, not the platform it runs on.
- If
The output settles an outcome on its own, such as rejecting a container on an inspection line.
ThenTreat it as critical: a static, deterministic model, full acceptance testing on independent data and change control from first deployment.
No human check stands between a wrong output and the product, which is the case draft Annex 22 addresses4.
- If
The model feeds a decision a trained operator makes, such as ranking deviations by likely root cause.
ThenWrite the operator's responsibility into the intended use and monitor operators like any manual step.
The draft allows reduced testing here only when the human role is defined and monitored4.
- If
A large language model drafts text inside a GxP process, such as a proposed deviation description.
ThenKeep it to non-critical work, have a qualified person confirm each output, and validate access, audit trail and how approved text enters the record.
The draft excludes generative AI from critical applications and expects human review elsewhere4.
- If
The model generates evidence for a submission, such as selecting manufacturing conditions.
ThenPlan a credibility assessment for its context of use and consider early agency discussion.
FDA's draft guidance covers AI informing decisions on safety, effectiveness or quality7.
A validation sequence for one model, from intended use to release
Describe the intended use and input space
State what the model decides, the process step it sits in and the full range of inputs, including rare variants. Divide inputs into subgroups that matter, such as defect type or site.
Assess risk feature by feature
Ask what happens to product, patient or data if each function fails and how anyone would notice. Spend scripted testing where failure is harmful and hard to detect.
Fix metrics and acceptance criteria
For a classifier, typically sensitivity, specificity and a confusion matrix per subgroup. Measure the current process first, because the model must do at least as well.
Build and lock the test set
Cover every subgroup, verify labels through independent experts or laboratory results, and store the set where developers cannot reach it, logging each access.
Run the approved test plan
Execute as written, log per-prediction confidence and the features behind key decisions, and justify every deviation before release.
Release under change and configuration control
Freeze model version, threshold and preprocessing, monitor performance and input drift, and route low-confidence outputs to a person as “undecided”.
The evidence file to keep for every validated model
When supplier, cloud and model-vendor evidence can be relied on
Annex 11 already expects supplier assessment and formal agreements with service providers2. For AI, that assessment must reach the model: who trained it, on which data and how versions are released. Vendor test results can reduce your testing when they cover your intended use, but they never replace testing on your own products and conditions.
Hosted foundation models may change without notice, which conflicts with configuration control, so pin versions contractually or self-host. ColdAI's life sciences delivery designs systems against GLP, GCP, GMP and Part 11 requirements with validation documentation8, and its enterprise AI practice runs inference inside client infrastructure9.
A visual inspection model on a fill-finish line, from risk assessment to release
Questions and answers
Can large language models be used in GxP work at all?
Yes, outside critical applications. Draft Annex 22 says generative AI and large language models should not be used where outputs directly affect patient safety, product quality or data integrity, and that qualified staff should confirm their outputs elsewhere4. EMA has said it is still considering consultation feedback that suggested support for enabling generative AI, and has sought expert input on controls such as guardrails10, so watch for the final text.
Do we validate the AI model or the computerized system around it?
Both, as one system. Inspectors look at the computerized system, its intended use and the process it supports, while the model needs its own test evidence. Draft Annex 22 places the model, the system and the process under change control together4. The validation plan covers infrastructure and interfaces as usual and adds a model-specific test plan, test set and monitoring.
Which changes to an AI model trigger revalidation?
Assess every change to the model, the system or the process, including changes to physical objects used as input, as draft Annex 22 notes4. Retraining creates a new version and normally needs acceptance testing against the locked criteria; threshold, preprocessing and camera or lighting changes usually do too. An infrastructure move may need only evidence that outputs are unchanged.
Where does FDA's draft guidance on AI for regulatory decision-making fit?
It covers AI producing information for decisions about a drug's safety, effectiveness or quality across nonclinical, clinical, postmarketing and manufacturing phases7. Its seven-step credibility framework sits alongside GMP validation: a manufacturing model may need both a validated system and a documented credibility assessment. Discovery and purely operational uses fall outside it.
Sources
- 21 CFR Part 11 — Electronic Records; Electronic Signatures — Legal Information Institute, Cornell Law School · checked 10 October 2026
- EudraLex Volume 4 — Good Manufacturing Practice guidelines, including Annex 11: Computerised Systems — European Commission · checked 10 October 2026
- Stakeholders' consultation on EudraLex Volume 4: Chapter 4, Annex 11 and new Annex 22 — European Commission · checked 10 October 2026
- Annex 22: Artificial Intelligence (consultation draft) — European Commission · checked 10 October 2026
- Computer Software Assurance for Production and Quality System Software; final guidance, notice of availability — Federal Register (FDA) · checked 10 October 2026
- ISPE GAMP 5 Guide: A Risk-Based Approach to Compliant GxP Computerized Systems (Second Edition) — ISPE · checked 10 October 2026
- Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products (draft guidance) — U.S. Food and Drug Administration · checked 10 October 2026
- Life Sciences: GxP-compliant architecture and delivery process — ColdAI
- Enterprise AI: on-premise and private-cloud deployments — ColdAI
- Good manufacturing practice: multistakeholder workshop on AI guidance development (Annex 22) — European Medicines Agency · checked 10 October 2026