This is the fact check of the episode as published. Each factual claim was extracted from the script and verified by an AI agent that saw only the claim and the primary document (the PDFs in the corpus). Verdicts: SUPPORTED (the document says it), PARTIAL (supported with a difference, noted), NOT-IN-CORPUS (the source is not among the primary documents on disk), NOT-CHECKABLE (an estimate or a characterisation). Opinions voiced by the hosts are listed but not verified. This report is itself AI output and can be wrong; corrections: jack@jackprior.ai.
These statements stand, but the primary documents on disk do not themselves confirm them: they come from the landscape reference (a working index whose rows are checked against the web), are estimates, or are hedged in the episode as such.
| Line | Speaker | Claim | Verdict | Why it stands | Evidence |
| 7 | SARAH | The ICH Quality Implementation Working Group's Points to Consider for Q8, Q9 and Q10 dates from 2011. | PARTIAL | The copy on disk is the EMA February 2012 edition; 2011 rests on the reference number | Header: 'February 2012, EMA/CHMP/ICH/902964/2011' (EMA transmission edition) |
| 7 | SARAH | The Points to Consider is an ICH-endorsed implementation guide, final, with a section on models graded low, medium, high. | PARTIAL | 'ICH-endorsed' and 'final' are not stated in the text; the document calls itself 'not intended to be new guidelines' | s.1 lines 2-7: 'The ICH Quality Implementation Working Group (Q-IWG) has prepared Points to Consider … not intended to be new guidelines'; s.5.1 low/medium/high models |
| 9 | SARAH | Draft Annex 22 dates from July 2025. | NOT-CHECKABLE | Draft carries no date; landscape row web-verified | Draft text carries no date; landscape row (web-verified) gives the July 2025 consultation |
| 9 | SARAH | BioPhorum's AI risk guidance is dated May 2026 and was released in June 2026. | PARTIAL | Document footer May 2026; the June release date is landscape-sourced | Footer on every page: '©BioPhorum Operations Group Ltd \| May 2026'; landscape gives the release as 2 June 2026 |
| 22 | HOST | The Points to Consider has been adopted for fourteen years. | NOT-CHECKABLE | Arithmetic from the 2011 reference; not stated in the document | Arithmetic from the 2011 reference; the document states no adoption date |
| 57 | SARAH | The crosswalk maps ICH 2011 low impact to tiers 0-1, medium to tier 2, high to tier 3; FDA/M15 model risk low, low-medium, medium-high, high. | PARTIAL | Crosswalk puts ICH low-to-medium at T1, not low | Crosswalk §3: T0 'Low-impact', T1 'Low-to-Medium', T2 'Medium-impact', T3 'High-impact' |
| 81 | SARAH | The Points to Consider puts the model's impact grade in the regulatory submission. | PARTIAL | 5.4 scales the submission's level of detail to impact; it does not say the grade itself is stated in the submission | s.5.4 lines 401-402: 'The level of detail for describing a model in a regulatory submission is dependent on the impact'; s.5.1 'For the purposes of regulatory submissions…' |
| Line | Speaker | Claim | Verdict | Evidence |
| 9 | SARAH | BioPhorum's AI risk guidance is dated May 2026 and was released in June 2026. | PARTIAL | Footer on every page: '©BioPhorum Operations Group Ltd \| May 2026'; landscape gives the release as 2 June 2026 |
| 44 | SARAH | BioPhorum's guidance was written by eight member companies with six more contributing. | SUPPORTED | p.5 Authors: Bristol Myers Squibb, Eli Lilly, Incyte, Organon, Pfizer, Regeneron, UCB, Vertex (eight) plus BioPhorum; Contributors: six further companies |
| 44 | SARAH | BioPhorum proposes one harmonised GxP AI risk framework. | SUPPORTED | 8.0 (p.22): 'harmonized framework recommended by the BioPhorum AI risk guidance workstream'; 8.1 'harmonized GxP risk framework' |
| 44 | SARAH | BioPhorum section 8.3 assesses risk in three elements: scope (system level or individual feature), source and characteristics, and analysis (consequence and likelihood). | SUPPORTED | 8.3 (p.23): 'three interdependent assessment elements – scope of risk, source and characteristics of risk, and impact of risk'; Step 3 'Risk analysis (impact and likelihood)' |
| 44 | SARAH | BioPhorum rates decision consequence low for internal efficiency, moderate for data integrity, process or compliance impact, and high for patient safety, product quality or regulatory impact. | SUPPORTED | 8.5 Step 3a (p.24): 'Low – internal efficiency or operational impact • Moderate – data integrity, process outcome or compliance impact • High – patient safety, product quality or regulatory impact' |
| 46 | SARAH | BioPhorum section 8.5 says 'model influence is directly determined by model maturity and serves as a proxy for likelihood'. | SUPPORTED | 8.5 Step 3b (p.24): 'model influence is directly determined by model maturity and serves as a proxy for likelihood occurrence' |
| 46 | SARAH | BioPhorum's maturity has two dimensions: autonomy (human in the loop, human on the loop, no oversight) and adaptiveness (rules-based, static, dynamic). | SUPPORTED | 8.5 (p.24): Autonomy (HITL / HOTL / no human oversight) and Adaptiveness (rules-based / static / dynamic) |
| 46 | SARAH | BioPhorum's 3x3 maturity matrix rates static with human in the loop as low influence, static with no human oversight as high, and dynamic with no oversight as high. | SUPPORTED | 8.5 matrix (p.24): Static+HITL Low; Static+No oversight High; Dynamic+No oversight High |
| 46 | SARAH | BioPhorum combines influence and consequence into a 3x3 composite risk. | SUPPORTED | 8.5 Step 4 (p.25): composite risk 3x3 combining model influence (likelihood) and decision consequence (severity) |
| 48 | SARAH | BioPhorum defines influence as the propensity to deviate from intended behaviour, used where an FMEA would put likelihood. | SUPPORTED | 8.5 Step 3b: 'greater propensity to deviate from intended performance or behavior'; 'analogous to the likelihood dimension in traditional FMEA models' |
| 48 | SARAH | In BioPhorum's matrix a rules-based system with no human oversight is moderate influence. | SUPPORTED | 8.5 matrix: Rules-based + No human oversight = Moderate |
| 50 | SARAH | BioPhorum says independent decision-limiting controls (independent release testing, orthogonal verification, redundant in-process sampling, process interlocks) may constrain autonomy and lower the influence rating, provided they operate independently of the AI output. | SUPPORTED | 8.5 (p.25): 'independent release testing, orthogonal verification, redundant in-process sampling or automated process interlocks … provided they operate independently of the AI output' |
| 50 | SARAH | BioPhorum section 8.4 says model type (probabilistic vs deterministic) is deliberately not an independent risk dimension because those risks are 'already captured through the autonomy and adaptiveness dimensions that define model influence'. | SUPPORTED | 8.4 (p.23), verbatim: 'already captured through the autonomy and adaptiveness dimensions that define model influence' |
| 50 | SARAH | BioPhorum calls that position consistent with GAMP and FDA guidance and says model type should instead raise the validation evidence required. | SUPPORTED | 8.4 (p.23): 'This position is consistent with GAMP AI and FDA guidance … model type should inform the rigor of validation evidence required' |
| 75 | SARAH | BioPhorum's governance section says AI systems should not execute critical or regulated actions without a qualified human, and 'must not directly execute electronic signatures'. | SUPPORTED | 7.0 (p.21), verbatim |
| 78 | SARAH | BioPhorum says which axis drove the composite risk should decide which controls are added: monitoring and autonomy constraints if influence, governance and validation evidence if consequence. | SUPPORTED | 8.6 (p.26): influence-driven → 'performance monitoring, drift detection and constraints on autonomy or adaptiveness'; consequence-driven → 'stronger governance, validation evidence and oversight mechanisms' |
| Line | Speaker | Claim | Verdict | Evidence |
| 9 | SARAH | The EU AI Act is Regulation 2024/1689, in force since August 2024, with Article 6 and Annex III on high-risk. | SUPPORTED | OJ L 12.7.2024; Art. 113 entry into force twentieth day after publication; Art. 6 'Classification rules for high-risk AI systems'; Annex III |
| 39 | SARAH | AI Act Article 6 defines high-risk two ways: a safety component of a product, or itself a product, covered by Annex I harmonisation legislation and requiring third-party conformity assessment; or the use cases in Annex III. | SUPPORTED | AI Act Art. 6(1)(a)-(b), Art. 6(2) |
| 39 | SARAH | Medical devices are on the AI Act's Annex I list. | SUPPORTED | AI Act Annex I Section A entries 11 (Reg. 2017/745 medical devices) and 12 (Reg. 2017/746 IVDs) |
| 39 | SARAH | AI Act Annex III lists eight areas: biometrics, critical infrastructure, education, employment, essential services, law enforcement, migration and border control, and justice. | SUPPORTED | AI Act Annex III: eight numbered areas (1 Biometrics … 8 Administration of justice and democratic processes) |
| 39 | SARAH | Pharmaceutical manufacturing is not on Annex I or Annex III of the AI Act. | SUPPORTED | No Annex I entry (1-20) or Annex III area concerns medicinal products; 'pharmaceutical' does not occur in the regulation |
| Line | Speaker | Claim | Verdict | Evidence |
| 27 | SARAH | Draft Annex 22 section 1 (scope) applies to computerised systems in manufacturing where AI models are used in 'critical applications with direct impact on patient safety, product quality or data integrity'. | SUPPORTED | Annex 22 s.1 lines 5-7, verbatim |
| 29 | SARAH | Draft Annex 22 covers static models only (models that do not adapt during use) and deterministic models only (identical inputs give identical outputs). | SUPPORTED | Annex 22 s.1 lines 12-13, 16-17 |
| 29 | SARAH | Draft Annex 22 says dynamic models and probabilistic models are not covered and should not be used in critical GMP applications. | SUPPORTED | Annex 22 s.1 lines 13-15, 17-19 |
| 29 | SARAH | Draft Annex 22 says 'the document does not apply to Generative AI and Large Language Models, and such models should not be used in critical GMP applications'. | SUPPORTED | Annex 22 s.1 lines 20-21 (original includes '(LLM)') |
| 31 | SARAH | Draft Annex 22 has no graded influence axis; its human-in-the-loop clause is the nearest it comes to distinguishing an advisory model within critical applications. | SUPPORTED | Annex 22 has no graded influence axis; but 3.3 and 10.5 single out the model-as-input-to-human-decision case within critical applications and allow diminished testing for it |
| 33 | SARAH | Draft Annex 22 section 3.3 on human in the loop says where a model gives input to a decision made by a human operator and the testing effort has been diminished because of that, the intended use must describe the operator's responsibility, and the operator's training and performance should be monitored 'like any other manual process'. | SUPPORTED | Annex 22 3.3 lines 52-56: '…monitored like any other manual process' |
| 35 | SARAH | Draft Annex 22 section 3.1 says the intended use and the specific tasks the model assists or automates should be described in detail, including a comprehensive characterisation of the input data and their common and rare variations, called the input sample space. | SUPPORTED | Annex 22 3.1 lines 39-43: 'all common and rare variations; i.e. the input sample space' |
| 35 | SARAH | Draft Annex 22 makes a process subject matter expert responsible for the intended-use description and the acceptance criteria. | SUPPORTED | Annex 22 3.1 lines 43-46 (description) and 4.2 lines 64-67 (acceptance criteria): 'A process subject matter expert (SME) should be responsible' |
| 36 | SARAH | Draft Annex 22 section 4.3 ('no decrease') says acceptance criteria for a model should be at least as high as the performance of the process it replaces. | SUPPORTED | Annex 22 4.3 lines 68-70: 'No decrease. The acceptance criteria of a model, should be at least as high as the performance of the process it replaces.' |
| 73 | SARAH | Draft Annex 22 allows generative AI in non-critical applications with a human in the loop. | SUPPORTED | Annex 22 s.1 lines 21-25 (HITL condition for non-critical use) |
| 81 | SARAH | Draft Annex 22's personnel principle names process subject matter experts, quality, data scientists, IT and consultants. | SUPPORTED | Annex 22 2.1 Personnel lines 30-31: 'process subject matter experts (SMEs), QA, data scientists, IT, and consultants' |
| 86 | HOST | Draft Annex 22 is six pages long. | SUPPORTED | Footers 'Page 1 of 6' … 'Page 6 of 6' |
| Line | Speaker | Claim | Verdict | Evidence |
| 7 | SARAH | The ICH Quality Implementation Working Group's Points to Consider for Q8, Q9 and Q10 dates from 2011. | PARTIAL | Header: 'February 2012, EMA/CHMP/ICH/902964/2011' (EMA transmission edition) |
| 7 | SARAH | The Points to Consider is an ICH-endorsed implementation guide, final, with a section on models graded low, medium, high. | PARTIAL | s.1 lines 2-7: 'The ICH Quality Implementation Working Group (Q-IWG) has prepared Points to Consider … not intended to be new guidelines'; s.5.1 low/medium/high models |
| 22 | HOST | The Points to Consider has been adopted for fourteen years. | NOT-CHECKABLE | Arithmetic from the 2011 reference; the document states no adoption date |
| 81 | SARAH | The Points to Consider puts the model's impact grade in the regulatory submission. | PARTIAL | s.5.4 lines 401-402: 'The level of detail for describing a model in a regulatory submission is dependent on the impact'; s.5.1 'For the purposes of regulatory submissions…' |
| Line | Speaker | Claim | Verdict | Evidence |
| 7 | SARAH | The Points to Consider is about models generally, not AI. | SUPPORTED | No occurrence of 'AI', 'artificial', 'machine learning' or 'neural' in the document; s.5 'A model is a simplified representation of a system using mathematical terms' |
| 11 | SARAH | Section 5 of the Points to Consider defines a model as a simplified representation of a system using mathematical terms and says models can be used at every stage of development and manufacturing. | SUPPORTED | s.5 lines 273-275: 'A model is a simplified representation of a system using mathematical terms … can be utilised at every stage of development and manufacturing' |
| 11 | SARAH | Section 5.1 says that for regulatory submissions an important factor is the model's contribution in assuring the quality of the product, and the level of oversight should be commensurate with the level of risk associated with the use of the specific model. | SUPPORTED | s.5.1 lines 286-288, verbatim |
| 13 | SARAH | Low-impact models support development, with formulation optimisation as the example. | SUPPORTED | s.5.1 lines 289-291: 'typically used to support product and/or process development (e.g. formulation optimisation)' |
| 13 | SARAH | Medium-impact models 'can be useful in assuring quality of the product but are not the sole indicators of product quality'; examples most design space models and many in-process controls. | SUPPORTED | s.5.1 lines 292-294, verbatim |
| 13 | SARAH | High-impact models are those whose prediction is a significant indicator of quality of the product; examples a chemometric model for product assay and a surrogate model for dissolution. | SUPPORTED | s.5.1 lines 295-298, verbatim |
| 15 | SARAH | The Points to Consider says an MSPC model used for continuous process verification along with a traditional method for release testing would likely be medium impact. | SUPPORTED | s.5.1 lines 322-324: 'would likely be classified as a medium-impact model' |
| 15 | SARAH | The Points to Consider says the same MSPC model used to support a surrogate for release testing in a real-time release approach would likely be high impact. | SUPPORTED | s.5.1 lines 324-326: 'would likely be classified as a high-impact model' |
| 15 | SARAH | The Points to Consider says a calibration model behind a near-infrared method is high impact if the method is used for release testing. | SUPPORTED | s.5.1 lines 312-315: NIR calibration model; 'if the method is used for release testing, then the model will be high-impact' |
| 15 | SARAH | The Points to Consider gives a feed-forward model adjusting compression parameters from incoming material attributes as a medium-impact example. | SUPPORTED | s.5.1 lines 329-331: 'a feed forward model to adjust compression parameters … could be classified as a medium-impact model' |
| 17 | SARAH | Section 5.3 lists validation/verification elements appropriate for high-impact models: acceptance criteria tied to the model's purpose, prediction accuracy against the reference method, validation on an external data set from batches not used to build the model, parallel testing with the reference method at implementation and through the lifecycle, and verification at commercial scale. | SUPPORTED | s.5.3 lines 369-389: acceptance criteria; accuracy vs reference method; external data set; parallel testing at implementation and through lifecycle; verification at commercial scale |
| 17 | SARAH | Section 5.3 says the applicability of those elements for medium and low-impact models can be considered case by case. | SUPPORTED | s.5.3 lines 366-368: 'can be considered on a case-by-case basis' |
| 19 | SARAH | Section 5.2 lists nine model development steps, the first being defining the purpose of the model. | SUPPORTED | s.5.2 lines 335-357: steps 1–9; '1. Defining the purpose of the model' |
| 19 | SARAH | Step eight of section 5.2 is evaluating the effect of prediction uncertainty on product quality and reducing residual risk through the control strategy, applying to high and medium-impact models. | SUPPORTED | s.5.2 lines 353-356, step 8 ('In certain cases … if appropriate … this can apply to high-impact and medium-impact models') |
| 19 | SARAH | Step nine of section 5.2 is documenting the model and planning its verification and update through the lifecycle, with the level of documentation dependent on the impact of the model. | SUPPORTED | s.5.2 lines 357-360, step 9: 'The level of documentation would be dependent on the impact of the model' |
| 65 | SARAH | The Points to Consider's surrogate example calls for parallel testing against the reference method through the lifecycle. | SUPPORTED | s.5.3 lines 385-387: 'parallel testing with the reference method during the initial stage of model implementation and can be repeated throughout the lifecycle' (a high-impact element, not specific to the surrogate example) |
| 67 | SARAH | The Points to Consider says an MSPC model for CPV alongside traditional release testing is medium impact and as a real-time release surrogate is high impact. | SUPPORTED | s.5.1 lines 322-326 |
| Line | Speaker | Claim | Verdict | Evidence |
| 9 | SARAH | ICH Q9 revision 1 on quality risk management was finalised in January 2023. | SUPPORTED | Cover: 'Q9(R1) Final version Adopted on 18 January 2023'; Step 4 adoption 18 January 2023 |
| 21 | SARAH | Q9(R1) states as a primary principle that the level of effort, formality and documentation of the quality risk management process should be commensurate with the level of risk. | SUPPORTED | Q9(R1) Section 3, primary principle, verbatim |
| 21 | SARAH | Q9(R1)'s section on formality says formality is not binary but a continuum, and how much to apply depends on uncertainty, importance and complexity. | SUPPORTED | Q9(R1) 5.1: 'formality can be considered a continuum (or spectrum)'; factors Uncertainty, Importance, Complexity |
| 21 | SARAH | Q9(R1) says resource constraints should not be used to justify the use of lower levels of formality. | SUPPORTED | Q9(R1) 5.1, verbatim |
| Line | Speaker | Claim | Verdict | Evidence |
| 7 | SARAH | The landscape's crosswalk file lays seven gradings side by side. | SUPPORTED | Crosswalk §1 'The gradings, side by side' (seven rows) |
| 31 | SARAH | The landscape crosswalk labels the absence of a middle tier in Annex 22 as gap G1. | SUPPORTED | Crosswalk §5: 'G1 — Annex 22 has no middle tier' |
| 53 | SARAH | The landscape crosswalk defines four tiers (0 exploratory, 1 advisory, 2 contributing control, 3 determinative) on influence and consequence with Annex 22 criticality as anchor. | SUPPORTED | Crosswalk §3 rows T0–T3 |
| 57 | SARAH | The crosswalk maps ICH 2011 low impact to tiers 0-1, medium to tier 2, high to tier 3; FDA/M15 model risk low, low-medium, medium-high, high. | PARTIAL | Crosswalk §3: T0 'Low-impact', T1 'Low-to-Medium', T2 'Medium-impact', T3 'High-impact' |
| 61 | SARAH | The crosswalk lists regulatory impact as absent from the 2011 grading and from Annex 22, as a gap. | SUPPORTED | Crosswalk §5 'G4 — Regulatory-impact axis missing from GMP frameworks … absent from PtC and Annex 22' |
| 73 | SARAH | The landscape crosswalk names the agent-drafting-a-deviation case as the open criticality question. | SUPPORTED | Crosswalk §4 row 'Agentic deviation management … Critical — LLM/agentic, should not be used' |