This is the fact check of the episode as published. Each factual claim was extracted from the script and verified by an AI agent that saw only the claim and the primary document (the PDFs in the corpus). Verdicts: SUPPORTED (the document says it), PARTIAL (supported with a difference, noted), NOT-IN-CORPUS (the source is not among the primary documents on disk), NOT-CHECKABLE (an estimate or a characterisation). Opinions voiced by the hosts are listed but not verified. This report is itself AI output and can be wrong; corrections: jack@jackprior.ai.
These statements stand, but the primary documents on disk do not themselves confirm them: they come from the landscape reference (a working index whose rows are checked against the web), are estimates, or are hedged in the episode as such.
| Line | Speaker | Claim | Verdict | Why it stands | Evidence |
| 9 | SARAH | The ten Good Machine Learning Practice principles were issued by FDA, Health Canada and MHRA in 2021 and adopted internationally as IMDRF N88 in January 2025. | PARTIAL | The 2021 FDA/Health Canada/MHRA origin is landscape-sourced; N88 itself is confirmed | Cover: 'IMDRF/AIML WG/N88 FINAL: 2025', '27 January 2025'; the document never names FDA, Health Canada, MHRA or 2021 |
| 9 | SARAH | BioPhorum's June 2026 AI risk guidance is from eight companies. | PARTIAL | Document dated May 2026; eight author companies plus six contributors | p.5 Authors: eight member companies plus BioPhorum; Contributors: six more; footer 'May 2026' |
| 18 | SARAH | Annex 22 is six pages, a supplement to Annex 11, and after section one is almost entirely about evidence. | PARTIAL | 'almost entirely about evidence' is a characterisation; sections 2 and 10 are principles and operation | Footers 'Page 6 of 6'; Scope 'additional guidance to Annex 11'; sections 3–9 are testing/evidence, 2 and 10 are principles and operation |
| 50 | SARAH | The CSA guidance is from CDRH and CBER with CDER consulted; draft September 2022; final 24 September 2025; reissued 3 February 2026. | PARTIAL | The 2022 draft date is not in the Feb 2026 text (docket number only) | Footnote 1 (CDRH/CBER with CDER, OCP, OII consulted); title page dates; docket FDA-2022-D-0795 is the only trace of 2022 |
| 50 | SARAH | The landscape records CSA as the document most cited as the model for AI validation in manufacturing. | PARTIAL | Landscape says 'frequently cited', the script 'most cited' | Landscape CSA row: 'frequently cited as a model for AI validation in manufacturing' |
| Line | Speaker | Claim | Verdict | Evidence |
| 4 | SARAH | Annex 22's sections three to nine follow the scope gate in section one. | SUPPORTED | Document map: 1 Scope, 2 Principles, 3 Intended Use … 9 Confidence, 10 Operation |
| 18 | SARAH | Annex 22 is six pages, a supplement to Annex 11, and after section one is almost entirely about evidence. | PARTIAL | Footers 'Page 6 of 6'; Scope 'additional guidance to Annex 11'; sections 3–9 are testing/evidence, 2 and 10 are principles and operation |
| 20 | SARAH | Annex 22 section 3: intended use and specific tasks described in detail based on in-depth knowledge of the process; comprehensive characterisation of input data and all common and rare variations, 'the input sample space'; limitations and possible erroneous or biased inputs identified; a process SME responsible for adequacy; documented and approved before acceptance testing starts. | SUPPORTED | 3.1 (pp.2-3), verbatim ('erroneous and biased inputs') |
| 22 | SARAH | Annex 22 3.2 divides the sample space into subgroups: decision output, process baseline characteristics (site, equipment), material or product, and task (types and severity of defects). | SUPPORTED | 3.2 (p.3), prefaced 'Where applicable' |
| 22 | SARAH | Annex 22 requires the test set to include the subgroups, allows acceptance criteria to differ by subgroup, and applies the size requirement to each subgroup. | SUPPORTED | 5.1 'stratified, include all subgroups'; 4.2 'may differ for specific subgroups'; 5.2 'any of its subgroups, should be sufficient in size' |
| 24 | SARAH | Annex 22 section 4: test metrics defined for the intended use; for a classifier examples are confusion matrix, sensitivity, specificity, accuracy, precision or F1; criteria may differ by subgroup; process SME responsible; approved before testing. | SUPPORTED | 4.1 metrics; 4.2 subgroup criteria, SME, approval before testing |
| 24 | SARAH | Annex 22 4.3 'No decrease': acceptance criteria 'at least as high as the performance of the process it replaces'; 'this implies that the performance should be known for the process which is to be replaced'; points to the revised Annex 11. | SUPPORTED | 4.3 (p.3), verbatim; cites 'Annex 11 2.7' |
| 27 | SARAH | Annex 22 section 5: test data representative of and expanding the full sample space, stratified, including all subgroups and rare variations, criteria and rationale documented; size sufficient for the whole set and each subgroup with adequate statistical confidence; labelling verified to 'a very high degree of correctness' (independent verification by multiple experts, validated equipment, or laboratory tests); pre-processing pre-specified; exclusions documented and justified. | SUPPORTED | 5.1–5.5 (pp.3-4) |
| 27 | SARAH | Annex 22 5.6: generating test data or labels, for example by generative AI, 'is not recommended and any use hereof should be fully justified'. | SUPPORTED | 5.6 (p.4): 'Generation of test data or labels, e.g. by means of generative AI, is not recommended and any use hereof should be fully justified' |
| 29 | SARAH | Annex 22 section 7: test should show the model is 'generalising well' including detection of over- or underfitting; approved test plan with intended use, metrics and criteria, reference to test data, test script, calculation method, process SME involved; deviations, failures or omission to use all test data documented, investigated and justified; everything retained including test data, physical objects, access-control and audit-trail records. | SUPPORTED | 7.1–7.4 (pp.4-5) |
| 32 | SARAH | Annex 22 section 6 'Test Data Independency'; 6.1 technical or procedural controls so test data are not used in development, training or validation; two routes (capture after training, or split before training). | SUPPORTED | 6, 6.1 (p.4): 'technical and/or procedural controls … capturing test data only after completion of training and validation, or by splitting' |
| 32 | SARAH | Annex 22 6.2: 'essential that employees involved in the development and training of the model have never had access to the test data'; access control and audit trail; 'there should be no copies of test data outside this repository'. | SUPPORTED | 6.2 (p.4), verbatim (conditional on the split route) |
| 34 | SARAH | Annex 22 6.3: record which data were used for testing, when, and how many times; 6.4: physical objects used for the final test must not previously have been used to train or validate; 6.5 Staff independency: controls preventing staff with test-data access from training or validating, else pair work under the '4-eyes principle'. | SUPPORTED | 6.3, 6.4 ('unless features are independent'), 6.5 '(4-eyes principle)' |
| 39 | SARAH | Annex 22 8.1 Feature attribution: during testing of models in critical GMP applications, systems should capture and record 'the features in the test data that have contributed to a particular classification or decision' (rejection as example); SHAP, LIME or heat maps should highlight key factors; 8.2 Feature justification: review of those features part of approving test results. | SUPPORTED | 8.1, 8.2 (p.5), verbatim ('Where applicable' for SHAP/LIME/heat maps) |
| 41 | SARAH | Annex 22 9.1: log the confidence score for each prediction where applicable; 9.2: appropriate threshold, and if confidence is very low consider whether the model should 'flag the outcome as undecided, rather than making potentially unreliable predictions or classifications'. | SUPPORTED | 9.1, 9.2 (p.5), verbatim |
| 55 | HOST | Annex 22 clause 2.3 sizes every activity to the risk to patient, product and data. | SUPPORTED | 2.3 (p.2): 'implemented based on the risk to patient safety, product quality and data integrity' |
| 68 | SARAH | Annex 22 2.2: documentation for these activities should be available to and reviewed by the regulated user irrespective of whether the model was trained, validated and tested in-house or by a supplier. | SUPPORTED | 2.2 (p.2): 'irrespective of whether a model is trained, validated and tested in-house or whether it is provided by a supplier or service provider' |
| 68 | SARAH | Annex 22 6.1 offers capturing the test set after training as a route to independence. | SUPPORTED | 6.1 (p.4) |
| 75 | SARAH | Annex 22 5.5 requires excluded test data to be documented. | SUPPORTED | 5.5 (p.4): 'documented and fully justified' |
| 77 | SARAH | Annex 22 says for non-critical use with a qualified human owning every output the principles may be considered where applicable. | SUPPORTED | Scope (p.2): the sentence is specific to generative AI/LLMs in non-critical use, which is the agent's case: '…the principles described in this document may be considered where applicable' |
| 84 | HOST | Annex 22 clause 2.1 puts MSAT, QA and the data scientist together at algorithm selection. | SUPPORTED | 2.1 (p.2): 'process subject matter experts (SMEs), QA, data scientists, IT, and consultants' during 'algorithm selection, and model training, validation, testing and operation' ('MSAT' is the script's mapping of process SMEs) |
| Line | Speaker | Claim | Verdict | Evidence |
| 4 | SARAH | FDA's draft step four is the credibility assessment plan and step six the report. | SUPPORTED | IV.A.4 lines 262-263 'credibility assessment plans'; IV.A.6 lines 486-487 'credibility assessment report' |
| 7 | SARAH | FDA's draft is from CDER with CBER, CDRH and other centres, January 2025, draft, with manufacturing named in its scope. | SUPPORTED | Footnote 1; cover 'DRAFT GUIDANCE … January 2025'; footnote 10 'drug product life cycle includes … manufacturing phases' |
| 12 | SARAH | FDA's scope section lists what scales with model risk (level of oversight, stringency of assessments and acceptance criteria, risk mitigation, extent of documentation) and says all should be 'commensurate with the AI model risk and tailored to the specific COU'. | SUPPORTED | Sec. II lines 53-61, verbatim |
| 12 | SARAH | FDA step four says performance acceptance criteria should be more stringent, and described in more detail, for high-risk models than low-risk ones; for certain low-risk models FDA may ask for minimal information. | SUPPORTED | Step 4 lines 290-291 and 298-299 |
| 14 | SARAH | FDA gives worked examples for steps one to three but not step four, because appropriate activities vary with the nuances of a specific programme. | SUPPORTED | IV.A lines 138-148 |
| 14 | SARAH | Step four has two halves: 4.a describes the model and its development (model, data, training); 4.b describes model evaluation on test data. Step five executes; step six documents results and deviations in a credibility assessment report; step seven decides adequacy for the COU. | SUPPORTED | 4.a line 293 (i model, ii data, iii training); 4.b lines 406-409; steps 5–7 lines 472-500 |
| 16 | SARAH | The credibility assessment report may be a self-contained document in a submission or meeting package, or 'held and made available to FDA on request, for example during an inspection'. | SUPPORTED | IV.A.6 lines 492-495 |
| 16 | SARAH | A footnote says that for uses outside the established meeting routes sponsors may complete all seven steps without seeking early engagement. | SUPPORTED | Footnote 25: for 'certain uses of AI … outside of contexts with established meeting options' (e.g. postmarketing pharmacovigilance) 'sponsors may choose to complete all the steps … without seeking early engagement' |
| 36 | SARAH | FDA step four: test data 'should be independent of the development data and should not be shown to the algorithm during training'; sponsor should specify how independence was achieved, e.g. 'data acquired using different batches or products'; overlapping use explained and justified; reference method described with a summary of its performance. | SUPPORTED | 4.b lines 410-434 |
| 44 | SARAH | FDA lists metrics (AUROC, sensitivity, specificity, predictive values, precision, F1) and says 'all performance estimates should be provided with confidence intervals'. | SUPPORTED | 4.a.iii lines 381-386: 'area under the receiver operating characteristic (ROC) curve … All performance estimates should be provided with confidence intervals' |
| 44 | SARAH | FDA asks sponsors to specify whether a pre-trained model was used and if so 'specify the dataset that was used for pre-training and how the pre-trained model was developed and/or obtained'. | SUPPORTED | 4.a.iii lines 393-396, verbatim |
| 46 | SARAH | FDA 4.b: if the COU involves a human in the loop, 'ensure that the evaluation methods consider the performance of the human-AI team, rather than just the performance of the model in isolation'. | SUPPORTED | 4.b lines 446-449, verbatim |
| 48 | SARAH | FDA 4.a asks for 'the quality assurance and control procedures of computer software (including its toolboxes and packages) and how version changes were tracked'; 4.b asks for procedures for code verification. | SUPPORTED | 4.a.iii lines 403-404; 4.b line 468, verbatim |
| 56 | SARAH | FDA's 2025 draft footnotes say question of interest, context of use and model risk were informed by ASME V&V 40 and point to the 2023 guidance for decision consequence. | SUPPORTED | Footnote 13 (ASME V&V40 sections 2, 3, 4); footnote 22 (November 2023 device guidance) |
| 62 | SARAH | FDA's example: Drug B parenteral multidose vial, fill volume a CQA, AI visual system on every vial, release testing on a sample per batch; high consequence, low influence, medium model risk. | SUPPORTED | IV.A.1 lines 166-169; IV.A.3 lines 252-253; sample at lines 204-205 |
| Line | Speaker | Claim | Verdict | Evidence |
| 9 | SARAH | FDA's CSA guidance for production and quality management system software is device-side, final September 2025 and reissued 3 February 2026. | SUPPORTED | Title page: CDRH/CBER; 'issued September 24, 2025'; 'Document issued on February 3, 2026' |
| 50 | SARAH | CSA is written for software in device production or a quality management system. | SUPPORTED | Sec. I: 'used as part of medical device production or the quality management system' |
| 50 | SARAH | CSA defines computer software assurance as 'a risk-based approach for establishing and maintaining confidence that software is fit for its intended use' and says 'the burden of validation is no more than necessary to address the risk'. | SUPPORTED | Section V (p.5), both quotes verbatim |
| 52 | SARAH | CSA's four steps: identify intended use; determine whether a failure poses high process risk (a quality problem that foreseeably compromises safety); choose assurance activities commensurate (unscripted testing incl. scenario, error-guessing, exploratory alongside scripted; unscripted may be better suited even for high-risk features; leverage vendor validation, other process controls, and monitoring data); establish the record (intended use, risk analysis, what was tested, issues, conclusion of acceptability, who/when, approval). | SUPPORTED | V.A sub-steps (1), (2), (4), (6) (the guidance has six, incl. (3) software changes and (5) additional considerations); high process risk definition; 'unscripted testing may be better suited … even for high process risk features' |
| 52 | SARAH | CSA names AI and machine learning tools, bots and cloud in its scope. | SUPPORTED | V.A (p.6): 'automation tools (e.g., BOTS or automatic workflows), data analytic tools, artificial intelligence/machine learning tools, and cloud computing' |
| 54 | SARAH | CSA: documentation 'need not include more evidence than necessary to show the software performs as intended for the risk identified'; FDA recommends 'system logs, audit trails, and other data generated and maintained by the software' rather than paper or screenshots. | SUPPORTED | V.A(6) lines 970-981, verbatim |