The European Commission has published the first version of a GMP guideline written specifically for artificial intelligence. Annex 22 to EudraLex Volume 4 sets out what manufacturers will be expected to do when an AI or machine-learning system touches the production of an active substance or a medicinal product. It is short, six pages, and it is not yet in force. It is also the clearest signal so far of what a regulator will ask for when AI enters a GMP process.
What changed
On 7 July 2025 the Commission opened a consultation on three documents at once: a revised Chapter 4 on documentation, a revised Annex 11 on computerised systems, and a wholly new Annex 22 on artificial intelligence. The drafts were prepared by the EMA GMDP Inspectors’ Working Group in cooperation with PIC/S, which means the expectations are intended to align across PIC/S’s participating authorities rather than the EU alone.
The comment period closed on 7 October 2025. As of August 2026 the annex has not been adopted: the EudraLex Volume 4 index still carries no Annex 22, and Annex 11 still shows its January 2011 revision. Neither the Commission nor EMA has announced an adoption date, and EMA has reopened the draft’s central scope question, which makes any predicted date speculation. Nothing here is law today. That is precisely why it is worth reading today.
Who it applies to
The instrument is Annex 22 of EudraLex Volume 4, the EU’s GMP guidance. It applies to AI and machine-learning models used in the manufacture of active substances and medicinal products, in critical applications with direct impact on patient safety, product quality, or data integrity. It does not cover every use of AI in a pharmaceutical company. It covers the uses that sit inside the regulated process.
Read it alongside the two documents it travelled with. Annex 11 governs the computerised systems an AI model runs on. Chapter 4 governs the records that prove what the system did. Annex 22 is the AI-specific layer on top of both.
What the draft excludes
The scope section does the draft’s most consequential work, and it does it in three steps. The document applies to static models; dynamic models, which continuously and automatically learn and adapt their performance during use, are not covered and should not be used in critical GMP applications (draft lines 12 to 15). It applies to models with deterministic output; models with probabilistic output, which given identical inputs might not produce identical outputs, are likewise not covered and should not be used in critical GMP applications (lines 16 to 19). Following those two, the draft states that it does not apply to generative AI and large language models, and that such models should not be used in critical GMP applications. In non-critical GMP applications, the ones with no direct impact on patient safety, product quality, or data integrity, personnel with adequate qualification and training should always be responsible for ensuring the outputs are suitable for the intended use: a human in the loop, with the document’s principles considered where applicable (lines 20 to 25).
These are three distinct exclusions in the draft’s own text: dynamic models, probabilistic-output models, and generative AI with LLMs as a named class. EMA’s later summary compresses them into a single statement; the draft’s wording, not the summary, is what a manufacturer would be held to. For anyone building agentic or LLM-based workflows near a GMP process, this is the paragraph of the draft that matters most. It is also exactly the paragraph the regulator has now reopened.
EMA has reopened the generative AI question
The exclusion did not survive the consultation unchallenged. In EMA’s own words: “A stakeholder consultation that EMA carried out in 2025 on its draft Annex 22 guidance suggested support for potentially enabling the use of technologies such as generative AI (GenAI) or large language models (LLMs) in medicines manufacturing.” The same page describes the draft’s position in the past tense: the draft “had indicated” that such models should not be used in critical GMP applications.
On 30 June and 1 July 2026 EMA held a multistakeholder workshop in Amsterdam: an open first day where experts from industry associations, academia, and regulatory authorities presented opinions and evidence, and a closed second day where the Annex 22 drafting group reviewed the contributions. The drafting group, together with EMA’s Quality Innovation Group, is seeking expert input on control and mitigation measures such as guardrails, as part of a proposed risk-based approach for manufacturers who want to use these technologies in GMP applications. EMA expects the workshop to produce a report. As of 12 August 2026 no report has been published and no revised draft has appeared.
Two things follow. First, while the drafting group is still working the scope question, any adoption date is a guess; that is why this page no longer carries one. Second, the direction of travel is visible. The question has moved from whether generative AI belongs in GMP at all to what controls would have to surround it.
Why it matters for AI workflows
Most discussion of AI in regulated work argues about whether the model is good enough. Annex 22 moves the question. It does not ask whether your model is clever. It asks whether you can show what it was for, how you tested it and with what data, who reviews its output, and how you would notice if it drifted.
The draft sets expectations in a recognisable shape. Define the model’s intended use before anything else. Establish the performance metrics that decide whether the model is acceptable for that use. Show that the test data is representative, documented, and demonstrably independent of the data the model was trained on, so the evaluation means something (sections 5 and 6). Validate the system. Then keep watching it: change control, performance monitoring, and, where the model feeds a human decision, a defined human-review point with records kept.
If you are putting an agent or a model anywhere near a GMP process, that list is the bar you will be measured against. For static, deterministic models it is the clearest statement of expectations a regulator has produced. For generative AI and LLMs the current draft would keep you out of critical applications entirely and behind a qualified human everywhere else, and the regulator is now consulting on the guardrails that could change that. Either way, the list is written down before you have to meet it.
The preflight implication
Every requirement in that list is a thing you can check before a workflow runs, not after an inspector asks. That is the whole argument for a preflight.
- Intended use declared. The workflow states what the model is for, in scope terms, before it acts.
- Test-data independence shown. The data that proves the model works is representative, documented, and demonstrably separate from what the model was trained on (sections 5 and 6).
- Validation evidence present. The performance metrics and the test results exist and are linked to the intended use.
- Human-review gate defined where one applies. Where a model feeds a human decision and testing has been reduced, the operator’s responsibility is described and records are kept, possibly a review of every output (sections 3.3 and 10.5). For generative AI in non-critical use, a qualified human responsible for the outputs is the standing expectation (lines 20 to 25).
- Ongoing monitoring in place. There is a defined way to notice drift and a change-control path when the model changes.
A preflight check asserts each of these on the ground, every time, before the work is in the air. It does not depend on how confident the model sounded. Annex 22 is, in effect, a regulator writing down the checklist. The work left for a manufacturer is to run that checklist before the process does, and to keep the evidence that it was run.
What you can do now
You cannot comply with a guideline that has not been adopted. You can do something more useful than wait. Map the AI-touching steps in your GMP workflows against the five expectations above and find the gaps now, while the cost of finding them is a note rather than a deviation. Build the evidence trail (intended use, test-data independence, validation, the human-review gate, monitoring) so that when Annex 22 is adopted you are reading your own records rather than starting them.
If your workflow is built on generative AI, the honest position is narrower: the scope question is open, and neither the exclusion nor its removal is safe to assume. What is safe to assume is the shape of whatever lands. The draft’s own text keeps a qualified human responsible for generative outputs, and EMA’s framing of the reopened question is about the controls that would surround such models. Each points at the same buildable things: defined controls around the model and a named human gate on its output, with evidence that both were in place. That part you can build, and check, today. The expensive version of this work is the one that begins after the rule is in force.
Sources
- European Commission, DG Health and Food Safety. Stakeholders’ Consultation on EudraLex Volume 4: Chapter 4, Annex 11 and new Annex 22. Consultation opened 7 July 2025, closed 7 October 2025. health.ec.europa.eu
- European Commission. Draft Annex 22: Artificial Intelligence. Consultation document, July 2025. Scope exclusions at lines 12 to 25. Consultation guideline PDF
- European Medicines Agency. Good manufacturing practice: Multistakeholder workshop on expert contributions to artificial intelligence guidance development (Annex 22). Amsterdam, 30 June to 1 July 2026. ema.europa.eu
- EudraLex Volume 4, EU Guidelines for Good Manufacturing Practice for Medicinal Products for Human and Veterinary Use. Index checked 12 August 2026: no Annex 22 listed; Annex 11 unchanged since its January 2011 revision. health.ec.europa.eu
Status note: Annex 22 is a draft. The comment period has closed. It is not in force, no adoption date has been announced, and EMA is reconsidering the draft’s treatment of generative AI. Dates and status above are current as of 12 August 2026 and will be updated here as the instrument moves.
Corpus anchor: this change touches eu-gmp-annex-11 (Computerised Systems) and pics-pi-011-3 (PIC/S Good Practices for Computerised Systems), both already in the Preclari corpus. Preclari checks against these today.