AI in Biopharma Manufacturing: Are We There Yet?

Fact check — Evidence: What Do I Have to Show?

Claims were extracted from the episode script and verified against the primary documents by AI agents that saw only the claim and the document. This report is itself AI output; it can be wrong. Corrections: jack@jackprior.ai.

Fact check — Evidence: What Do I Have to Show?

This is the fact check of the episode as published. Each factual claim was extracted from the script and verified by an AI agent that saw only the claim and the primary document (the PDFs in the corpus). Verdicts: SUPPORTED (the document says it), PARTIAL (supported with a difference, noted), NOT-IN-CORPUS (the source is not among the primary documents on disk), NOT-CHECKABLE (an estimate or a characterisation). Opinions voiced by the hosts are listed but not verified. This report is itself AI output and can be wrong; corrections: jack@jackprior.ai.

Episode as publishedClaims
Claims checked58
Supported by the primary text52
Confirmed by official web sources4
Resting on secondary sources or hedged as such1
Corrected before release1
Open0
Host opinions (not verified)24

Claims resting on secondary sources

These statements stand, but the primary documents on disk do not themselves confirm them: they come from the landscape reference (a working index whose rows are checked against the web), are estimates, or are hedged in the episode as such.

LineSpeakerClaimVerdictWhy it standsEvidence
9SARAHThe ten Good Machine Learning Practice principles were issued by FDA, Health Canada and MHRA in 2021 and adopted internationally as IMDRF N88 in January 2025.PARTIALThe 2021 FDA/Health Canada/MHRA origin is landscape-sourced; N88 itself is confirmedCover: 'IMDRF/AIML WG/N88 FINAL: 2025', '27 January 2025'; the document never names FDA, Health Canada, MHRA or 2021
9SARAHBioPhorum's June 2026 AI risk guidance is from eight companies.PARTIALDocument dated May 2026; eight author companies plus six contributorsp.5 Authors: eight member companies plus BioPhorum; Contributors: six more; footer 'May 2026'
18SARAHAnnex 22 is six pages, a supplement to Annex 11, and after section one is almost entirely about evidence.PARTIAL'almost entirely about evidence' is a characterisation; sections 2 and 10 are principles and operationFooters 'Page 6 of 6'; Scope 'additional guidance to Annex 11'; sections 3–9 are testing/evidence, 2 and 10 are principles and operation
50SARAHThe CSA guidance is from CDRH and CBER with CDER consulted; draft September 2022; final 24 September 2025; reissued 3 February 2026.PARTIALThe 2022 draft date is not in the Feb 2026 text (docket number only)Footnote 1 (CDRH/CBER with CDER, OCP, OII consulted); title page dates; docket FDA-2022-D-0795 is the only trace of 2022
50SARAHThe landscape records CSA as the document most cited as the model for AI validation in manufacturing.PARTIALLandscape says 'frequently cited', the script 'most cited'Landscape CSA row: 'frequently cited as a model for AI validation in manufacturing'

All claims, by document

BioPhorum-2026-06-AI-Risk-Guidance-Harmonizing-Frameworks.pdf

LineSpeakerClaimVerdictEvidence
9SARAHBioPhorum's June 2026 AI risk guidance is from eight companies.PARTIALp.5 Authors: eight member companies plus BioPhorum; Contributors: six more; footer 'May 2026'
77SARAHBioPhorum's June 2026 worked example of a deviation module proposes back-testing against closed deviations and CAPAs, cause-evidence explainability, mandatory investigator justification, hallucination flags and citation requirements.SUPPORTEDBioPhorum Example 2 (p.29): feature C 'back-testing versus closed deviations/CAPAs; cause-evidence explainability; mandatory investigator justification'; feature A 'hallucination flags and citation requirements'
81SARAHBioPhorum's example grades deviation drafting as low consequence (operational efficiency) with a static locked model and human approval.SUPPORTEDBioPhorum Example 2 feature A (p.29): 'Low (operational efficiency) … Low (static, locked; human approval) … Composite: Low'

EU-2025-07-Annex-22-AI-Draft.pdf

LineSpeakerClaimVerdictEvidence
4SARAHAnnex 22's sections three to nine follow the scope gate in section one.SUPPORTEDDocument map: 1 Scope, 2 Principles, 3 Intended Use … 9 Confidence, 10 Operation
18SARAHAnnex 22 is six pages, a supplement to Annex 11, and after section one is almost entirely about evidence.PARTIALFooters 'Page 6 of 6'; Scope 'additional guidance to Annex 11'; sections 3–9 are testing/evidence, 2 and 10 are principles and operation
20SARAHAnnex 22 section 3: intended use and specific tasks described in detail based on in-depth knowledge of the process; comprehensive characterisation of input data and all common and rare variations, 'the input sample space'; limitations and possible erroneous or biased inputs identified; a process SME responsible for adequacy; documented and approved before acceptance testing starts.SUPPORTED3.1 (pp.2-3), verbatim ('erroneous and biased inputs')
22SARAHAnnex 22 3.2 divides the sample space into subgroups: decision output, process baseline characteristics (site, equipment), material or product, and task (types and severity of defects).SUPPORTED3.2 (p.3), prefaced 'Where applicable'
22SARAHAnnex 22 requires the test set to include the subgroups, allows acceptance criteria to differ by subgroup, and applies the size requirement to each subgroup.SUPPORTED5.1 'stratified, include all subgroups'; 4.2 'may differ for specific subgroups'; 5.2 'any of its subgroups, should be sufficient in size'
24SARAHAnnex 22 section 4: test metrics defined for the intended use; for a classifier examples are confusion matrix, sensitivity, specificity, accuracy, precision or F1; criteria may differ by subgroup; process SME responsible; approved before testing.SUPPORTED4.1 metrics; 4.2 subgroup criteria, SME, approval before testing
24SARAHAnnex 22 4.3 'No decrease': acceptance criteria 'at least as high as the performance of the process it replaces'; 'this implies that the performance should be known for the process which is to be replaced'; points to the revised Annex 11.SUPPORTED4.3 (p.3), verbatim; cites 'Annex 11 2.7'
27SARAHAnnex 22 section 5: test data representative of and expanding the full sample space, stratified, including all subgroups and rare variations, criteria and rationale documented; size sufficient for the whole set and each subgroup with adequate statistical confidence; labelling verified to 'a very high degree of correctness' (independent verification by multiple experts, validated equipment, or laboratory tests); pre-processing pre-specified; exclusions documented and justified.SUPPORTED5.1–5.5 (pp.3-4)
27SARAHAnnex 22 5.6: generating test data or labels, for example by generative AI, 'is not recommended and any use hereof should be fully justified'.SUPPORTED5.6 (p.4): 'Generation of test data or labels, e.g. by means of generative AI, is not recommended and any use hereof should be fully justified'
29SARAHAnnex 22 section 7: test should show the model is 'generalising well' including detection of over- or underfitting; approved test plan with intended use, metrics and criteria, reference to test data, test script, calculation method, process SME involved; deviations, failures or omission to use all test data documented, investigated and justified; everything retained including test data, physical objects, access-control and audit-trail records.SUPPORTED7.1–7.4 (pp.4-5)
32SARAHAnnex 22 section 6 'Test Data Independency'; 6.1 technical or procedural controls so test data are not used in development, training or validation; two routes (capture after training, or split before training).SUPPORTED6, 6.1 (p.4): 'technical and/or procedural controls … capturing test data only after completion of training and validation, or by splitting'
32SARAHAnnex 22 6.2: 'essential that employees involved in the development and training of the model have never had access to the test data'; access control and audit trail; 'there should be no copies of test data outside this repository'.SUPPORTED6.2 (p.4), verbatim (conditional on the split route)
34SARAHAnnex 22 6.3: record which data were used for testing, when, and how many times; 6.4: physical objects used for the final test must not previously have been used to train or validate; 6.5 Staff independency: controls preventing staff with test-data access from training or validating, else pair work under the '4-eyes principle'.SUPPORTED6.3, 6.4 ('unless features are independent'), 6.5 '(4-eyes principle)'
39SARAHAnnex 22 8.1 Feature attribution: during testing of models in critical GMP applications, systems should capture and record 'the features in the test data that have contributed to a particular classification or decision' (rejection as example); SHAP, LIME or heat maps should highlight key factors; 8.2 Feature justification: review of those features part of approving test results.SUPPORTED8.1, 8.2 (p.5), verbatim ('Where applicable' for SHAP/LIME/heat maps)
41SARAHAnnex 22 9.1: log the confidence score for each prediction where applicable; 9.2: appropriate threshold, and if confidence is very low consider whether the model should 'flag the outcome as undecided, rather than making potentially unreliable predictions or classifications'.SUPPORTED9.1, 9.2 (p.5), verbatim
55HOSTAnnex 22 clause 2.3 sizes every activity to the risk to patient, product and data.SUPPORTED2.3 (p.2): 'implemented based on the risk to patient safety, product quality and data integrity'
68SARAHAnnex 22 2.2: documentation for these activities should be available to and reviewed by the regulated user irrespective of whether the model was trained, validated and tested in-house or by a supplier.SUPPORTED2.2 (p.2): 'irrespective of whether a model is trained, validated and tested in-house or whether it is provided by a supplier or service provider'
68SARAHAnnex 22 6.1 offers capturing the test set after training as a route to independence.SUPPORTED6.1 (p.4)
75SARAHAnnex 22 5.5 requires excluded test data to be documented.SUPPORTED5.5 (p.4): 'documented and fully justified'
77SARAHAnnex 22 says for non-critical use with a qualified human owning every output the principles may be considered where applicable.SUPPORTEDScope (p.2): the sentence is specific to generative AI/LLMs in non-critical use, which is the agent's case: '…the principles described in this document may be considered where applicable'
84HOSTAnnex 22 clause 2.1 puts MSAT, QA and the data scientist together at algorithm selection.SUPPORTED2.1 (p.2): 'process subject matter experts (SMEs), QA, data scientists, IT, and consultants' during 'algorithm selection, and model training, validation, testing and operation' ('MSAT' is the script's mapping of process SMEs)

FDA-2023-11-Credibility-Computational-Modeling.pdf

LineSpeakerClaimVerdictEvidence
9SARAHFDA's November 2023 device guidance on credibility of computational modelling wrote down the credibility vocabulary.SUPPORTEDTitle page 'November 17, 2023'; Section IV Definitions (four terms reprinted from ASME V&V 40-2018; 'adequacy assessment' is FDA's own)
56SARAHFDA's 2023 credibility guidance says it covers first-principles models and not machine learning.SUPPORTEDSection III Scope (p.7): 'not intended to apply to standalone statistical or data-driven models such as … machine learning or artificial intelligence-based models'
56SARAHFDA 2023 defines credibility as trust, established through the collection of evidence, in a model's predictive capability for a context of use; verification concerns the code and the calculation.SUPPORTEDSection IV (p.9) credibility definition verbatim; 'Code verification and calculation verification are two elements of verification'
58SARAHFDA 2023 defines validation as 'the process of determining the degree to which a model or a simulation is an accurate representation of the real world', requiring comparison against data independent of the data used to create the model (so calibration is not validation).SUPPORTEDSection IV (p.10) verbatim; VI.A: 'model calibration … is not considered validation'
58SARAHFDA 2023 defines applicability as 'the relevance of the validation activities to support the use of the computational model for a context of use'; credibility factors each have a gradation of rigour and a credibility goal chosen by model risk; adequacy assessment asks whether the evidence together is sufficient for the risk.SUPPORTEDSection IV (p.8) verbatim; VI.C steps 5.2–5.3 gradation and credibility goal; adequacy assessment definition

FDA-2025-01-AI-Regulatory-Decision-Making-Draft.pdf

LineSpeakerClaimVerdictEvidence
4SARAHFDA's draft step four is the credibility assessment plan and step six the report.SUPPORTEDIV.A.4 lines 262-263 'credibility assessment plans'; IV.A.6 lines 486-487 'credibility assessment report'
7SARAHFDA's draft is from CDER with CBER, CDRH and other centres, January 2025, draft, with manufacturing named in its scope.SUPPORTEDFootnote 1; cover 'DRAFT GUIDANCE … January 2025'; footnote 10 'drug product life cycle includes … manufacturing phases'
12SARAHFDA's scope section lists what scales with model risk (level of oversight, stringency of assessments and acceptance criteria, risk mitigation, extent of documentation) and says all should be 'commensurate with the AI model risk and tailored to the specific COU'.SUPPORTEDSec. II lines 53-61, verbatim
12SARAHFDA step four says performance acceptance criteria should be more stringent, and described in more detail, for high-risk models than low-risk ones; for certain low-risk models FDA may ask for minimal information.SUPPORTEDStep 4 lines 290-291 and 298-299
14SARAHFDA gives worked examples for steps one to three but not step four, because appropriate activities vary with the nuances of a specific programme.SUPPORTEDIV.A lines 138-148
14SARAHStep four has two halves: 4.a describes the model and its development (model, data, training); 4.b describes model evaluation on test data. Step five executes; step six documents results and deviations in a credibility assessment report; step seven decides adequacy for the COU.SUPPORTED4.a line 293 (i model, ii data, iii training); 4.b lines 406-409; steps 5–7 lines 472-500
16SARAHThe credibility assessment report may be a self-contained document in a submission or meeting package, or 'held and made available to FDA on request, for example during an inspection'.SUPPORTEDIV.A.6 lines 492-495
16SARAHA footnote says that for uses outside the established meeting routes sponsors may complete all seven steps without seeking early engagement.SUPPORTEDFootnote 25: for 'certain uses of AI … outside of contexts with established meeting options' (e.g. postmarketing pharmacovigilance) 'sponsors may choose to complete all the steps … without seeking early engagement'
36SARAHFDA step four: test data 'should be independent of the development data and should not be shown to the algorithm during training'; sponsor should specify how independence was achieved, e.g. 'data acquired using different batches or products'; overlapping use explained and justified; reference method described with a summary of its performance.SUPPORTED4.b lines 410-434
44SARAHFDA lists metrics (AUROC, sensitivity, specificity, predictive values, precision, F1) and says 'all performance estimates should be provided with confidence intervals'.SUPPORTED4.a.iii lines 381-386: 'area under the receiver operating characteristic (ROC) curve … All performance estimates should be provided with confidence intervals'
44SARAHFDA asks sponsors to specify whether a pre-trained model was used and if so 'specify the dataset that was used for pre-training and how the pre-trained model was developed and/or obtained'.SUPPORTED4.a.iii lines 393-396, verbatim
46SARAHFDA 4.b: if the COU involves a human in the loop, 'ensure that the evaluation methods consider the performance of the human-AI team, rather than just the performance of the model in isolation'.SUPPORTED4.b lines 446-449, verbatim
48SARAHFDA 4.a asks for 'the quality assurance and control procedures of computer software (including its toolboxes and packages) and how version changes were tracked'; 4.b asks for procedures for code verification.SUPPORTED4.a.iii lines 403-404; 4.b line 468, verbatim
56SARAHFDA's 2025 draft footnotes say question of interest, context of use and model risk were informed by ASME V&V 40 and point to the 2023 guidance for decision consequence.SUPPORTEDFootnote 13 (ASME V&V40 sections 2, 3, 4); footnote 22 (November 2023 device guidance)
62SARAHFDA's example: Drug B parenteral multidose vial, fill volume a CQA, AI visual system on every vial, release testing on a sample per batch; high consequence, low influence, medium model risk.SUPPORTEDIV.A.1 lines 166-169; IV.A.3 lines 252-253; sample at lines 204-205

FDA-2026-02-CSA-Production-Quality-System-Software.pdf

LineSpeakerClaimVerdictEvidence
9SARAHFDA's CSA guidance for production and quality management system software is device-side, final September 2025 and reissued 3 February 2026.SUPPORTEDTitle page: CDRH/CBER; 'issued September 24, 2025'; 'Document issued on February 3, 2026'
50SARAHCSA is written for software in device production or a quality management system.SUPPORTEDSec. I: 'used as part of medical device production or the quality management system'
50SARAHCSA defines computer software assurance as 'a risk-based approach for establishing and maintaining confidence that software is fit for its intended use' and says 'the burden of validation is no more than necessary to address the risk'.SUPPORTEDSection V (p.5), both quotes verbatim
52SARAHCSA's four steps: identify intended use; determine whether a failure poses high process risk (a quality problem that foreseeably compromises safety); choose assurance activities commensurate (unscripted testing incl. scenario, error-guessing, exploratory alongside scripted; unscripted may be better suited even for high-risk features; leverage vendor validation, other process controls, and monitoring data); establish the record (intended use, risk analysis, what was tested, issues, conclusion of acceptability, who/when, approval).SUPPORTEDV.A sub-steps (1), (2), (4), (6) (the guidance has six, incl. (3) software changes and (5) additional considerations); high process risk definition; 'unscripted testing may be better suited … even for high process risk features'
52SARAHCSA names AI and machine learning tools, bots and cloud in its scope.SUPPORTEDV.A (p.6): 'automation tools (e.g., BOTS or automatic workflows), data analytic tools, artificial intelligence/machine learning tools, and cloud computing'
54SARAHCSA: documentation 'need not include more evidence than necessary to show the software performs as intended for the risk identified'; FDA recommends 'system logs, audit trails, and other data generated and maintained by the software' rather than paper or screenshots.SUPPORTEDV.A(6) lines 970-981, verbatim

FDA-2026-02-CSA-Production-Quality-System-Software.pdf; FDA-2022-09-FR-Notice-CSA-Draft-Guidance.pdf

LineSpeakerClaimVerdictEvidence
50SARAHThe CSA guidance is from CDRH and CBER with CDER consulted; draft September 2022; final 24 September 2025; reissued 3 February 2026.PARTIALFootnote 1 (CDRH/CBER with CDER, OCP, OII consulted); title page dates; docket FDA-2022-D-0795 is the only trace of 2022

ICH-2011-Q8Q9Q10-Points-to-Consider.pdf

LineSpeakerClaimVerdictEvidence
25HOSTThe Points to Consider describes parallel testing against the reference method as a comparator alongside a model.SUPPORTEDPtC 5.3 lines 385-387 'parallel testing with the reference method'; 5.1 lines 322-323 'along with a traditional method for release testing'; 'confirmatory' does not appear

IMDRF-2025-N88-GMLP-Guiding-Principles.pdf

LineSpeakerClaimVerdictEvidence
44SARAHThe IMDRF GMLP principles say generative systems may employ foundation models not under the manufacturer's provenance.SUPPORTEDIMDRF Introduction (p.4): 'may employ foundation models that are not under the provenance of the medical device manufacturers'
46SARAHIMDRF principle seven concerns the human-AI team and lists human factors: user skills, expertise, understanding of the model's limitations, and potential for over-reliance.SUPPORTEDIMDRF principle 7 (pp.7-8): 'user skills, user expertise, user understanding of the model outputs and limitations, potential for overreliance'
60SARAHIMDRF N88 is final, dated 27 January 2025, and lists ten GMLP principles: multidisciplinary expertise; good software engineering and security; representative datasets; training sets independent of test sets with 'the extent of external validation is proportionate to risk'; reference standards fit for purpose; model choice tailored to data and intended use; human-AI team assessed; testing under real conditions; clear information to users; deployed models monitored with retraining risks managed.SUPPORTEDCover 'FINAL: 2025', '27 January 2025'; ten principle headings; principle 4 'The extent of external validation is proportionate to risk'

IMDRF-2025-N88-GMLP-Guiding-Principles.pdf; FDA-2025-12-GMLP-Guiding-Principles-Page.pdf

LineSpeakerClaimVerdictEvidence
9SARAHThe ten Good Machine Learning Practice principles were issued by FDA, Health Canada and MHRA in 2021 and adopted internationally as IMDRF N88 in January 2025.PARTIALCover: 'IMDRF/AIML WG/N88 FINAL: 2025', '27 January 2025'; the document never names FDA, Health Canada, MHRA or 2021

AI-CMC-Regulatory-Landscape.md (landscape self-description)

LineSpeakerClaimVerdictEvidence
50SARAHThe landscape records CSA as the document most cited as the model for AI validation in manufacturing.PARTIALLandscape CSA row: 'frequently cited as a model for AI validation in manufacturing'

Not among the primary documents on disk

LineSpeakerClaimVerdictEvidence
7SARAHAnnex 22 was out for comment July to October 2025 with roughly 1,300 comments, text unchanged after the EMA workshop in summer 2026; drafted by EMA's inspectors working group and consulted with PIC/S. [consultation, comments, workshop, provenance: landscape]NOT-IN-CORPUSPartly on disk now. EMA minutes, 4 Feb 2026 (EMA/40804/2026), §3, p. 2: 'GMP Annex 22 on AI in manufacturing, which following public consultation received ~1,300 public comments and is undergoing revision. The final document is expected to be published by the end of the year.' — confirms the ~1,300 comments and a final expected by end-2026. EMA event page (EMA-2026-06-Annex22-Workshop-Page.pdf, captured 6 Sep 2026): 'EMA's Good Manufacturing Practice (GMP) / Good Distribution Practice (GDP) Inspectors Working Group is organising a two-day workshop to help shape a risk-based approach to the use of generative artificial intelligence (AI) in medicines manufacturing.' … 'The draft Annex 22 had indicated that dynamic, adaptive and probabilistic models - such as GenAI or LLMs - should not be used in critical GMP applications. EMA is still considering the implications of the stakeholder consultation results.' — confirms the 30 Jun–1 Jul 2026 workshop, EMA's Inspectors Working Group as owner, and that the July 2025 draft still stood with EMA 'still considering' the consultation. The rest: landscape Annex 22 row.

Host opinions (listed, not verified)