This is the fact check of the episode as published. Each factual claim was extracted from the script and verified by an AI agent that saw only the claim and the primary document (the PDFs in the corpus). Verdicts: SUPPORTED (the document says it), PARTIAL (supported with a difference, noted), NOT-IN-CORPUS (the source is not among the primary documents on disk), NOT-CHECKABLE (an estimate or a characterisation). Opinions voiced by the hosts are listed but not verified. This report is itself AI output and can be wrong; corrections: jack@jackprior.ai.
These statements stand, but the primary documents on disk do not themselves confirm them: they come from the landscape reference (a working index whose rows are checked against the web), are estimates, or are hedged in the episode as such.
| Line | Speaker | Claim | Verdict | Evidence |
| 6 | SARAH | BioPhorum's June 2026 guidance is the only one of the six documents that defines its oversight terms. | SUPPORTED | Glossary (p.30): 'Human approval (HITL)', 'Human-on-the-loop (HOTL)', 'Human supervision'; none of the other five documents carries such definitions |
| 12 | SARAH | BioPhorum's glossary, citing ISO, defines human-in-the-loop (labelled human approval: a person must review, approve, modify or block the output before any consequential action), human-on-the-loop (system operates autonomously, a person supervises and can intervene), and human supervision (a person oversees but does not approve every decision). | SUPPORTED | Glossary p.30: 'Human approval (HITL) … (Source: ISO)'; 'Human-on-the-loop (HOTL) … (ISO/IEC 2289 intent)'; 'Human supervision — … does not approve every decision' (no source cited for the third) |
| 12 | SARAH | BioPhorum's risk matrix uses three columns: human-in-the-loop, human-on-the-loop, no human oversight. | SUPPORTED | 8.5 Step 3b matrix headers: HITL, HOTL, No human oversight |
| 14 | SARAH | BioPhorum: autonomy and adaptiveness are the two axes of model maturity; maturity is what BioPhorum calls model influence, its likelihood term. | SUPPORTED | 8.5 Step 3b: 'model influence is directly determined by model maturity and serves as a proxy for likelihood occurrence' |
| 14 | SARAH | BioPhorum matrix: static with HITL low; static with HOTL moderate; static with no oversight high; dynamic moderate with HITL and high in the other two columns. | SUPPORTED | 8.5 matrix: Static Low/Moderate/High; Dynamic Moderate/High/High |
| 16 | SARAH | BioPhorum says human oversight is one mechanism by which autonomy may be limited, and other decision-limiting controls (independent release testing, orthogonal verification, redundant in-process sampling, automated interlocks) should count when assigning autonomy. | SUPPORTED | 8.5 (p.25): 'Human oversight (approval or supervision) is one mechanism by which autonomy may be limited; however, other systematic decision-limiting controls … should be considered for purposes of assigning autonomy' |
| 48 | SARAH | BioPhorum governance: human oversight plays a dual role and can function as both a risk control and a risk factor; HITL mechanisms are essential but may introduce variability, bias or informal workarounds if roles, competencies and decision boundaries are not clearly defined. | SUPPORTED | 7.0 (p.21), verbatim |
| 48 | SARAH | BioPhorum, citing Annex 22 on adaptive models: reliance on human intervention alone is insufficient as a long-term control because system behaviour may evolve outside the assumptions of the original validation state. | SUPPORTED | 7.0 (p.21): 'as recognized in EU GMP Annex 22 (Draft), reliance on human intervention alone is insufficient as a long-term control' |
| 50 | SARAH | BioPhorum: effective AI governance must explicitly define when and how human oversight functions as a validated control, and when it constitutes an additional source of risk requiring training, procedural controls and ongoing monitoring. | SUPPORTED | 7.0 (p.21), verbatim ('therefore' omitted) |
| 52 | SARAH | BioPhorum's appendix example grades a QMS deviation module feature by feature: for ranked root-cause suggestions it lists overreliance by users as a risk source and mandatory investigator justification as a control; for the proposed CAPA it rates consequence high and influence moderate with human approval, composite high. | SUPPORTED | Example 2 (p.29): row C source 'overreliance by users', control 'mandatory investigator justification'; row D consequence High, influence Moderate (human approval), composite High |
| 58 | SARAH | BioPhorum says 'a qualified human'. | SUPPORTED | 7.0 (p.21): 'without intervention from a qualified human (i.e. HITL)' |
| 60 | SARAH | BioPhorum names an AI owner, identified during development, who stays accountable for fitness for use throughout the lifecycle. | SUPPORTED | 7.0 (p.21): 'An AI owner should be identified during the development … accountability for AI fitness for use remains with the AI owner throughout the lifecycle' |
| 60 | SARAH | BioPhorum: AI systems must not directly execute electronic signatures, and should not be permitted to execute critical or regulated actions without intervention from a qualified human. | SUPPORTED | 7.0 (p.21): 'AI systems should not be permitted to execute critical or regulated actions or transactions without intervention from a qualified human (i.e. HITL). Additionally, AI systems must not directly execute electronic signatures.' |
| 75 | HOST | BioPhorum's example asks for citation requirements and back-testing against closed deviations. | SUPPORTED | Example 2 (p.29): row A 'citation requirements'; row C 'back-testing versus closed deviations/CAPAs' |
| Line | Speaker | Claim | Verdict | Evidence |
| 8 | SARAH | EMA's reflection paper was issued by CHMP with CVMP; drafted July 2023, consulted to the end of 2023, adopted final on 9 September 2024. | SUPPORTED | Cover: draft agreed July 2023; consultation 19 July – 31 December 2023; final adopted by CHMP 9 September 2024 (CVMP 11 September) |
| 8 | SARAH | EMA's manufacturing section is one paragraph; its oversight content sits in the deployment, governance and ethics sections. | SUPPORTED | 2.3.6 is a single paragraph; oversight content in 2.5.6, 2.6 and 2.8 |
| 40 | SARAH | EMA 2.5.6 model deployment: for all models, especially those with no human-in-the-loop, a system risk management plan should be developed defining likely risks of failure modes of the algorithm, covering consequences of incorrect predictions or classifications, monitoring and mitigation or correction approaches, how to trigger suspension or decommissioning and how to carry it out. | SUPPORTED | 2.5.6 (p.11), near-verbatim |
| 42 | SARAH | EMA 2.2 puts responsibility for ensuring algorithms, models, datasets and pipelines are fit for purpose with the sponsor, applicant, MAH or manufacturer. | SUPPORTED | 2.2 (p.4): 'clinical trial sponsor, marketing authorisation applicant/holder or manufacturer' |
| 42 | SARAH | EMA 2.6 governance: SOPs implementing GxP principles on data and algorithm governance should be extended to all data, models and algorithms used for AI where regulatory impact or patient risk is high. | SUPPORTED | 2.6 (p.11), verbatim |
| 42 | SARAH | EMA 2.8 lists seven ethics principles from the Commission's high-level expert group, the first being human agency and oversight, and says a human-centric approach should guide all development and deployment. | SUPPORTED | 2.8 (p.12): seven bulleted principles from the High-Level Expert Group on AI, first 'Human agency and oversight'; 'a human-centric approach should guide all development and deployment' |
| 44 | SARAH | EMA 2.3.6: AI in manufacturing from process design and scale-up to in-process control and batch release expected to increase; QRM principles with patient safety, data integrity and product quality; ICH Q8, Q9, Q10 'awaiting revision of current regulatory requirements and GMP standards'. | SUPPORTED | 2.3.6 (p.7); quoted phrase verbatim |
| 46 | SARAH | EMA's conclusion says efforts should be made in all organisations to 'reciprocally integrate data science competence with the respective fields within medicines development, manufacturing and pharmacovigilance'. | SUPPORTED | 3. Conclusion (p.12), verbatim |
| 47 | HOST | EMA's precision-medicine section asks for guidance prescribers can critically apprehend, and fall-back treatment strategies in cases of technical failure. | SUPPORTED | 2.3.4 Precision medicine (p.7): 'guidance that the prescribers can critically apprehend and to include fall-back treatment strategies in cases of technical failure' |
| 74 | SARAH | EMA warns generative models produce plausible but erroneous or incomplete output, written for product information. | SUPPORTED | 2.3.5 Product information, verbatim |
| Line | Speaker | Claim | Verdict | Evidence |
| 6 | SARAH | AI Act Article 14 is titled human oversight. | SUPPORTED | Article 14 'Human oversight' (OJ p.60) |
| 18 | SARAH | AI Act Regulation 2024/1689 in force since August 2024; Article 14 applies only to high-risk AI systems. | SUPPORTED | Art. 113 (entry into force 20 days after OJ 12.7.2024); every paragraph of Art. 14 addresses 'high-risk AI systems'; Art. 14 applies from 2 August 2026 (Art. 113) |
| 20 | SARAH | Article 14(1): high-risk systems designed, including with appropriate human-machine interface tools, so natural persons can effectively oversee them while in use. | SUPPORTED | Art. 14(1), verbatim substance |
| 20 | SARAH | Article 14(2): oversight aims to prevent or minimise risks in intended use or foreseeable misuse, in particular 'where such risks persist despite the application of other requirements'. | SUPPORTED | Art. 14(2): 'in particular where such risks persist despite the application of other requirements set out in this Section' |
| 20 | SARAH | Article 14(3): measures commensurate with the risks, level of autonomy and context of use, built in by the provider or implemented by the deployer. | SUPPORTED | Art. 14(3)(a)-(b) (measures identified by the provider; (b) implemented by the deployer) |
| 22 | SARAH | Article 14(4) lists five things overseers must be enabled to do: understand capacities and limitations and monitor operation (detect anomalies, dysfunctions, unexpected performance); remain aware of automation bias; correctly interpret the output; decide not to use, or disregard, override or reverse the output; intervene or interrupt through a stop button or similar. | SUPPORTED | Art. 14(4)(a)–(e), five points, in the order spoken; chapeau 'as appropriate and proportionate'; (e) 'stop button or a similar procedure' |
| 50 | SARAH | The AI Act names automation bias and requires overseers to remain aware of it. | SUPPORTED | Art. 14(4)(b) '(automation bias)' |
| 73 | HOST | AI Act Article 14(4) says the overseer must be able to disregard, override or reverse the output. | SUPPORTED | Art. 14(4)(d), verbatim |
| Line | Speaker | Claim | Verdict | Evidence |
| 24 | SARAH | Annex 22 uses the phrase human-in-the-loop in three places, and in the two operating clauses (3.3, 10.5) the word it uses for the human is operator. | SUPPORTED | Scope para (once, 'personnel … i.e. a human-in-the-loop (HITL)'); 3.3 (heading and text); 10.5 — three places, four occurrences; 3.3 and 10.5 say 'human operator', the scope paragraph says 'personnel' |
| 26 | SARAH | Annex 22 scope: generative AI/LLMs should not be used in critical GMP applications; in non-critical applications personnel with adequate qualification and training should always be responsible for ensuring outputs are suitable, glossed as a human-in-the-loop. | SUPPORTED | Scope lines 20-25, verbatim |
| 28 | SARAH | Annex 22 2.1 Personnel: close cooperation between process SMEs, QA, data scientists, IT and consultants during algorithm selection, training, validation, testing and operation, with adequate qualifications, defined responsibilities and appropriate level of access. | SUPPORTED | 2.1, verbatim |
| 30 | SARAH | Annex 22 3.3 Human-in-the-loop, in the intended-use section: where a model gives input to a decision by a human operator and testing effort has been diminished, the intended use should include the operator's responsibility; training and consistent performance monitored 'like any other manual process'. | SUPPORTED | 3.3 under '3. Intended Use', verbatim |
| 30 | SARAH | Annex 22 10.5: in the same situation records should be kept, and depending on criticality and level of testing this may imply a consistent review or test of every output, according to a procedure. | SUPPORTED | 10.5 Human review ('review and/or test of every output') |
| 56 | SARAH | Annex 22 3.1: intended use includes a comprehensive characterisation of the input sample space, a process SME responsible; 9.2: threshold and undecided flag on very low confidence; 10.3: system performance monitored against metrics; 10.4: monitor whether inputs remain within the sample space and intended use, with drift metrics. | SUPPORTED | 3.1, 9.2 ('should be considered whether the model should flag the outcome as undecided'), 10.3, 10.4 |