This is the fact check of the episode as published. Each factual claim was extracted from the script and verified by an AI agent that saw only the claim and the primary document (the PDFs in the corpus). Verdicts: SUPPORTED (the document says it), PARTIAL (supported with a difference, noted), NOT-IN-CORPUS (the source is not among the primary documents on disk), NOT-CHECKABLE (an estimate or a characterisation). Opinions voiced by the hosts are listed but not verified. This report is itself AI output and can be wrong; corrections: jack@jackprior.ai.
These statements stand, but the primary documents on disk do not themselves confirm them: they come from the landscape reference (a working index whose rows are checked against the web), are estimates, or are hedged in the episode as such.
| Line | Speaker | Claim | Verdict | Why it stands | Evidence |
| 5 | HOST | Draft Annex 22 is the rulebook's first binding-track text on AI. | NOT-CHECKABLE | 'first binding-track' is the landscape's characterisation | Landscape row: 'First binding-track GMP text on AI'; the draft says only 'new annex' |
| 7 | SARAH | Draft Annex 22 to the EU GMP guide is titled Artificial Intelligence and is an annex to EudraLex Volume 4. | PARTIAL | 'EudraLex Volume 4' does not appear in the draft; only the Annex 11 reference implies the GMP annex series | Line 1: 'Annex 22: Artificial Intelligence'; 'Reasons for changes: Not applicable (new annex)'. 'EudraLex' and 'Volume 4' do not appear; only the Annex 11 reference implies the GMP annex series |
| 7 | SARAH | Draft Annex 22 was published for consultation on 7 July 2025 and the consultation closed on 7 October 2025. | NOT-CHECKABLE | Draft carries no date; landscape row web-verified | Draft carries no date; landscape row (web-verified): consultation 7 Jul – 7 Oct 2025 |
| 7 | SARAH | Per the landscape, draft Annex 22 was prepared by the GMP and GDP Inspectors Working Group at EMA and consulted in parallel with PIC/S. | NOT-IN-CORPUS | Hedged in-script; the draft names no author | Landscape row: 'Drafted by EMA GMP/GDP Inspectors Working Group; co-published with PIC/S'; the draft names no author |
| 73 | HOST | The eight companies behind the BioPhorum guidance argued in June 2026 that model type should not be an exclusion. | PARTIAL | Position supported (8.4); the document is dated May 2026 and lists eight author companies plus six contributors | BioPhorum 8.4 (p.23): model type 'already captured through the autonomy and adaptiveness dimensions'; footer May 2026; eight author companies plus six contributors |
| 82 | SARAH | Draft Annex 22 is the first binding-track GMP text on AI, six pages. | NOT-CHECKABLE | 'first binding-track' is the landscape's characterisation; six pages confirmed | Landscape row; draft says 'new annex'; six pages confirmed |
| Line | Speaker | Claim | Verdict | Evidence |
| 9 | SARAH | EMA's reflection paper (final September 2024) addresses generative and incremental learning. | SUPPORTED | Cover 9 September 2024; 2.3.5 generative language models; 2.3.3.3 and 2.3.7 incremental learning |
| 50 | SARAH | EMA's product-information section says AI used for drafting, compiling, editing or reviewing product information should be used under close human supervision because 'generative language models are prone to include plausible but erroneous or incomplete output', with quality review to ensure text is factually and syntactically correct. | SUPPORTED | 2.3.5, verbatim |
| 50 | SARAH | EMA's integrity (data protection) section says large language models, often containing billions of parameters, are at particular risk of memorising training data. | SUPPORTED | 2.7: 'Large language models, often containing billions of parameters, are at particular risk of memorisation due to their size' |
| 50 | SARAH | EMA on manufacturing: model development and lifecycle management should follow quality risk management principles and ICH Q8, Q9 and Q10, 'awaiting revision of current regulatory requirements and GMP standards'. | SUPPORTED | 2.3.6, verbatim |
| 52 | SARAH | EMA on pivotal clinical trials: 'incremental learning approaches are not accepted', and any modification of the model during the trial requires a regulatory interaction. | SUPPORTED | 2.3.3.3 'Pivotal clinical trials': 'Incremental learning approaches are not accepted, and any modification of the model during the trial requires a regulatory interaction to amend the statistical analysis plan' |
| 52 | SARAH | EMA on pharmacovigilance: incremental learning can continuously enhance models for classifying adverse event reports, with the holder responsible for validating and monitoring. | SUPPORTED | 2.3.7: 'incremental learning can continuously enhance models for classification and severity scoring of adverse event reports … responsibility of the MAH to validate, monitor and document' |
| 52 | SARAH | EMA's glossary defines a frozen model as one where all parameters have been finally set, not allowing further adaption to new data, the same words as Annex 22's static definition. | SUPPORTED | Glossary: 'Frozen model: A model where all parameters have been finally set, not allowing further adaption to new data' |
| Line | Speaker | Claim | Verdict | Evidence |
| 7 | SARAH | Draft Annex 22 to the EU GMP guide is titled Artificial Intelligence and is an annex to EudraLex Volume 4. | PARTIAL | Line 1: 'Annex 22: Artificial Intelligence'; 'Reasons for changes: Not applicable (new annex)'. 'EudraLex' and 'Volume 4' do not appear; only the Annex 11 reference implies the GMP annex series |
| 7 | SARAH | Draft Annex 22 is six pages. | SUPPORTED | Footers 'Page 1 of 6' to 'Page 6 of 6' |
| 11 | SARAH | Annex 22 section 1 first paragraph: applies to computerised systems used in manufacturing medicinal products and active substances 'where Artificial Intelligence models are used in critical applications with direct impact on patient safety, product quality or data integrity'. | SUPPORTED | Scope lines 5-7, verbatim |
| 11 | SARAH | Annex 22 says it provides additional guidance to Annex 11 for computerised systems in which AI models are embedded. | SUPPORTED | Scope lines 7-8: 'provides additional guidance to Annex 11 for computerised systems in which AI models are embedded' |
| 13 | SARAH | Annex 22 second paragraph: machine learning models are models that obtained their functionality through training with data rather than being explicitly programmed; a model may consist of several individual models each automating a specific process step. | SUPPORTED | Scope lines 9-11: 'obtained their functionality through training with data, rather than being explicitly programmed. Models may consist of several individual models, each automating specific process steps' |
| 17 | SARAH | Annex 22 applies to static models, 'models that do not adapt their performance during use by incorporating new data'. | SUPPORTED | Scope lines 12-13, verbatim |
| 17 | SARAH | Annex 22: the use of dynamic models which continuously and automatically learn and adapt performance during use is not covered by this document and 'should not be used in critical GMP applications'. | SUPPORTED | Scope lines 13-15, verbatim |
| 19 | SARAH | Annex 22 uses 'should' throughout as its operative word. | SUPPORTED | 'should' occurs 62 times; whole-word 'must' and 'shall' occur nowhere |
| 19 | SARAH | Annex 22's glossary defines static as 'Frozen model: a model where all parameters have been finally set, not allowing further adaption to new data'. | SUPPORTED | Glossary lines 183-184, verbatim (spelling 'adaption') |
| 23 | SARAH | Annex 22 applies to models with a deterministic output 'which, when given identical inputs, provide identical outputs'; models with a probabilistic output are not covered and should not be used in critical GMP applications. | SUPPORTED | Scope lines 16-19 |
| 25 | SARAH | Annex 22's third scope paragraph opens with 'Following the above' and says 'the document does not apply to Generative AI and Large Language Models, and such models should not be used in critical GMP applications'. | SUPPORTED | Scope lines 20-21, verbatim; it is the fifth paragraph of the scope, the third of the static/deterministic/generative trio |
| 27 | SARAH | Annex 22: if generative models are used in non-critical GMP applications (those without direct impact on patient safety, product quality or data integrity), 'personnel with adequate qualification and training should always be responsible for ensuring that the outputs from such models are suitable for the intended use', a human-in-the-loop, and the principles in the document may be considered where applicable. | SUPPORTED | Scope lines 21-25, verbatim |
| 29 | SARAH | Annex 22 clause 3.3: where a model gives input to a decision made by a human operator and the testing effort has been diminished because of that, the intended use should include the operator's responsibility, and the operator's training and consistent performance should be monitored 'like any other manual process'. | SUPPORTED | 3.3 lines 52-56: '…monitored like any other manual process' |
| 29 | SARAH | Annex 22 clause 10.5 says records should be kept of the human review, and depending on criticality and the extent of testing this may mean review of every output according to a procedure. | SUPPORTED | 10.5 lines 159-163: 'records should be kept … may imply a consistent review and/or test of every output from the model, according to a procedure' |
| 58 | SARAH | Annex 22 clause 4.3 (no decrease) says the model's acceptance criteria should be at least as high as the performance of the process it replaces, and cites Annex 11 for how that performance is known. | SUPPORTED | 4.3 lines 68-70: 'should be at least as high as the performance of the process it replaces … (see Annex 11 2.7)' |
| 58 | SARAH | Annex 22 section 10 (operation) covers change control on the model, the system and the process before deployment; configuration control with measures to detect unauthorised change; performance monitoring against defined metrics; and monitoring whether inputs stay inside the model's sample space, with drift metrics. | SUPPORTED | 10.1–10.4 lines 144-158: change control; configuration control 'detect any unauthorised change'; performance monitoring; input sample space monitoring 'Metrics should be defined for monitoring any drift' |
| 60 | SARAH | Annex 22 covers intended use with an input sample space, acceptance criteria fixed before testing, independent test data, explainability, confidence, and operation clauses. | SUPPORTED | Sections 3, 4, 6, 8, 9, 10 |
| 64 | SARAH | Annex 22 clause 10.2 is configuration control; section 6 covers test data controls; section 4 fixes acceptance criteria; clause 10.1 is change control. | SUPPORTED | 10.2 line 150; 6 line 87; 4 line 57; 10.1 line 144 |
| 68 | SARAH | Annex 22 clause 10.1: a model change is evaluated for retest before going live, and a decision not to retest must be fully justified. | SUPPORTED | 10.1 lines 146-149: 'evaluated to determine if the model needs to be retested. Any decision not to conduct such retest should be fully justified' |
| 76 | SARAH | Annex 22's personnel clause puts process experts, quality, data scientists and IT together during algorithm selection, before training, validation and testing. | SUPPORTED | 2.1 lines 27-31: 'close cooperation … during algorithm selection, and model training, validation, testing and operation … process subject matter experts (SMEs), QA, data scientists, IT, and consultants' |
| 85 | HOST | Annex 22's evidence clauses are sections 3 to 10. | SUPPORTED | Document map sections 3–10 are the substantive requirement clauses; section 2 Principles also carries 'should' requirements |
| Line | Speaker | Claim | Verdict | Evidence |
| 9 | SARAH | FDA's device guidance on predetermined change control plans for AI-enabled device software was final in December 2024 and reissued in August 2025. | SUPPORTED | Title page: 'Document issued on August 18, 2025. Document originally issued on December 4, 2024.' |
| 41 | SARAH | The PCCP guidance for AI-enabled device software functions was originally issued 4 December 2024 and reissued 18 August 2025. | SUPPORTED | Title page dates, verbatim |
| 41 | SARAH | The PCCP guidance applies to device software the manufacturer intends to modify over time, including modifications implemented automatically by software ('continuous learning'), manually with human review and decision, or both. | SUPPORTED | Section III Scope (p.6): 'implemented automatically by software, also known as continuous learning … manually … or a combination of both' |
| 43 | SARAH | A PCCP has three components: a Description of Modifications, a Modification Protocol (verification and validation with pre-defined acceptance criteria), and an Impact Assessment of benefits and risks. | SUPPORTED | Section V.A (p.10): 'Description of Modifications, a Modification Protocol, and an Impact Assessment' |
| 43 | SARAH | An authorised PCCP becomes a technological characteristic of the authorised device, and modifications made within it do not trigger a new marketing submission. | SUPPORTED | Section IV (p.9): 'An authorized PCCP is a technological characteristic of the authorized device'; Section V: 'without triggering the need for a new marketing submission' |
| 43 | SARAH | The PCCP guidance says it applies to the device constituent of a combination product and not to the drug or biologic constituent. | SUPPORTED | Section III Scope (p.8): 'apply to the device constituent part of device-led combination products … do not apply to the drug or biologic constituent part' |
| Line | Speaker | Claim | Verdict | Evidence |
| 9 | SARAH | FDA's January 2025 draft has a scope carve-out and a section on lifecycle maintenance. | SUPPORTED | Sec. II line 47 (scope carve-out); Sec. IV.B line 510 'Life Cycle Maintenance of the Credibility of AI Model Outputs' |
| 33 | SARAH | FDA's January 2025 draft says in its introduction 'This guidance does not endorse the use of any specific AI approach or technique'. | SUPPORTED | Sec. I line 22, verbatim |
| 33 | SARAH | Model type never appears as a gate anywhere in the FDA draft. | SUPPORTED | 'model type', 'type of model' absent; architecture appears only as a step-4 description item (line 312) |
| 35 | SARAH | FDA section IV.B, lifecycle maintenance, is defined as managing changes to AI models whether incidentally or deliberately so the model stays fit for use over the product lifecycle, and says lifecycle maintenance is important for AI in the manufacturing phase specifically. | SUPPORTED | IV.B lines 513-515 and 524-526 |
| 35 | SARAH | FDA draft: AI-based models may be highly sensitive to changes in inputs because they are data-driven and can be 'self-evolving (i.e., capable of autonomously adapting without any human intervention)'; performance metrics should be monitored on an ongoing basis; 'sponsors should anticipate inherent, model-directed changes and the need to identify and evaluate those changes'. | SUPPORTED | IV.B lines 528-535, verbatim |
| 37 | SARAH | FDA draft: changes to the model or manufacturing changes that may affect it go through the manufacturer's change management system within the PQS, with newly available manufacturing data and model-directed changes as examples; performance-affecting changes should be reported to the Agency in accordance with regulatory requirements. | SUPPORTED | IV.B lines 539-550 |
| 37 | SARAH | FDA draft: detailed lifecycle plans (performance metrics, risk-based monitoring frequency, retesting triggers) sit in the site's PQS with a summary in the marketing application for product- or process-specific models. | SUPPORTED | IV.B lines 552-556 |
| 37 | SARAH | FDA draft: sponsors may propose model-related elements as established conditions under ICH Q12, with a plan to manage changes to them, so that some changes would not require submission before being made. | SUPPORTED | IV.B lines 559-567: 'which changes would not require submission to the Agency prior to making modifications' |
| 39 | SARAH | FDA step seven: one of the five outcomes when credibility is not sufficiently established is that the sponsor may change the modelling approach. | SUPPORTED | Step 7 lines 500-506: outcome (4) 'the sponsor may change the modeling approach' |
| 46 | SARAH | The words generative and language model do not appear in FDA's draft. | SUPPORTED | 'generative', 'language model', 'LLM': no matches |
| 46 | SARAH | FDA's scope carve-out at lines 47-51: does not address AI used 'for operational efficiencies' (internal workflows, resource allocation, drafting/writing a regulatory submission) that do not impact patient safety, drug quality, or the reliability of results from a nonclinical or clinical study. | SUPPORTED | Sec. II lines 47-50, verbatim (also excludes drug discovery) |
| 46 | SARAH | FDA adds that sponsors uncertain whether a use is in scope should engage early. | SUPPORTED | Lines 50-51: 'We encourage sponsors to engage with FDA early if they are uncertain about whether or not their use of AI is within the scope' |
| 48 | SARAH | FDA step four asks sponsors to specify whether a pre-trained model was used and, if so, which dataset it was pre-trained on and how the model was developed or obtained. | SUPPORTED | Step 4.a.iii lines 393-396, verbatim |
| 48 | SARAH | FDA step four says if the context of use involves a human in the loop, evaluate the human-AI team, not the model alone. | SUPPORTED | Step 4.b lines 446-449, verbatim |
| Line | Speaker | Claim | Verdict | Evidence |
| 9 | SARAH | NIST's generative AI profile is AI 600-1, from July 2024. | SUPPORTED | Title page: 'NIST AI 600-1 … Generative Artificial Intelligence Profile, July 2024' |
| 54 | SARAH | NIST AI 600-1 is a profile of the AI Risk Management Framework for generative AI, approved in July 2024, voluntary. | SUPPORTED | Section 1: 'cross-sectoral profile of and companion resource for the AI Risk Management Framework'; 'intended for voluntary use'; approved 07-25-2024 |
| 54 | SARAH | NIST AI 600-1 section 2 lists twelve risks unique to or exacerbated by generative AI. | SUPPORTED | Section 2 list items 1–12, subsections 2.1–2.12 |
| 54 | SARAH | NIST defines confabulation as 'the production of confidently stated but erroneous or false content', colloquially hallucinations, and says confabulation is especially important to monitor when generative AI is integrated into consequential decision-making. | SUPPORTED | Section 2 item 2 and 2.2, verbatim ('hallucinations or fabrications') |
| 54 | SARAH | NIST's Human-AI Configuration risk covers automation bias and over-reliance, humans perceiving generated content as higher quality than it is. | SUPPORTED | 2.7: 'may over-rely on GAI systems or may unjustifiably perceive GAI content to be of higher quality … automation bias' |
| 54 | SARAH | NIST's Information Integrity risk covers content that may not distinguish fact from inference or acknowledge its uncertainties. | SUPPORTED | Section 2 item 8: 'may not distinguish fact from opinion or fiction or acknowledge uncertainties'; 2.8 'distinguishes fact from fiction, opinion, and inference' |
| 54 | SARAH | NIST's Value Chain and Component Integration risk covers third-party components, pre-trained models and datasets not transparently vetted, making issues hard to attribute to a source. | SUPPORTED | 2.12: 'procured datasets, pre-trained models, and software libraries … not properly vetted … difficult to attribute issues in a system's behavior to any one of these sources' |