AI in Biopharma Manufacturing: Are We There Yet?

Fact check — Data: Are My Data Fit for Use?

Claims were extracted from the episode script and verified against the primary documents by AI agents that saw only the claim and the document. This report is itself AI output; it can be wrong. Corrections: jack@jackprior.ai.

Fact check — Data: Are My Data Fit for Use?

This is the fact check of the episode as published. Each factual claim was extracted from the script and verified by an AI agent that saw only the claim and the primary document (the PDFs in the corpus). Verdicts: SUPPORTED (the document says it), PARTIAL (supported with a difference, noted), NOT-IN-CORPUS (the source is not among the primary documents on disk), NOT-CHECKABLE (an estimate or a characterisation). Opinions voiced by the hosts are listed but not verified. This report is itself AI output and can be wrong; corrections: jack@jackprior.ai.

Episode as publishedClaims
Claims checked69
Supported by the primary text65
Confirmed by official web sources4
Resting on secondary sources or hedged as such0
Corrected before release0
Open0
Host opinions (not verified)16

Claims resting on secondary sources

These statements stand, but the primary documents on disk do not themselves confirm them: they come from the landscape reference (a working index whose rows are checked against the web), are estimates, or are hedged in the episode as such.

LineSpeakerClaimVerdictWhy it standsEvidence
4SARAHThe Annex 11 revision was consulted July to October 2025 in the same package as Annex 22 and the revised Chapter 4.NOT-CHECKABLELandscape-sourcedNeither draft carries a date; landscape row (web-verified)
6SARAHThe Annex 11 revision grows from about five pages to nineteen, in seventeen chapters: quality system, risk management, supplier management, qualification and validation, handling of data, identity and access, audit trails, electronic signatures, periodic review, security, backup, archiving.PARTIAL19 pages and 17 chapters confirmed; 'about five pages' for the current annex is landscape-sourced; the spoken list names 12 of the 17 chaptersFooter 'Page 19 of 19'; document map 1 Scope … 17 Archiving; the current five-page Annex 11 is not in the corpus
6SARAHFinal Annex 11 text expected around end 2026 and implementation around early 2027 (estimates).NOT-CHECKABLEHedged in-scriptLandscape watchlist estimate
41SARAHAI Act high-risk obligations are phased in through 2027 and 2028 (hedged 'as I understand it').NOT-CHECKABLEHedged 'as I understand it'Landscape §3.3 (Digital Omnibus); the AI Act text on disk still says 'apply from 2 August 2026'

All claims, by document

EMA-2024-09-AI-Medicinal-Product-Lifecycle-Reflection-Paper.pdf

LineSpeakerClaimVerdictEvidence
10SARAHEMA's reflection paper has a section on data acquisition.SUPPORTED2.5.1 'Data acquisition and augmentation'
47SARAHEMA 2.5.1: models are data-driven and vulnerable to bias; acquire a balanced and sufficiently large training set for the intended context of use; sources, acquisition and processing (cleaning, transformation, imputation, annotation, normalisation, augmentation) 'documented in a detailed and fully traceable manner in line with GxP requirements'; exploratory analysis records relevance, representativeness, interpolation/extrapolation assumptions, class imbalances; augmentation may expand the training set; remaining limitations presented in model documentation.SUPPORTED2.5.1 (p.8), verbatim phrases
49SARAHEMA 2.5.2: an early train-test split before any normalisation is strongly encouraged because the risk of data leakage, 'often unintentional or even unconscious', cannot be completely excluded.SUPPORTED2.5.2 (p.9): '…strongly encouraged. Even so, the risk of direct or indirect (often unintentional or even unconscious) data leakage cannot be completely excluded'
51SARAHEMA says augmentation and synthetic data may in some cases be useful for expanding the training set.SUPPORTED2.5.1: 'synthetic data of other modalities may in some cases be useful for expanding the training dataset'
53SARAHEMA's model deployment section says the data acquisition hardware, software and data transformation pipeline at inference should be in line with pre-defined specifications.SUPPORTED2.5.6 Model deployment (p.10)

EMA-2026-03-GMDP-IWG-Work-Plan-2026-2028.pdf

LineSpeakerClaimVerdictEvidence
6SARAHFinal Annex 11 text expected around end 2026 and implementation around early 2027 (estimates).NOT-CHECKABLELandscape watchlist estimate

EU-2024-AI-Act-Reg-2024-1689.pdf

LineSpeakerClaimVerdictEvidence
10SARAHThe EU AI Act's Article 10 is on data governance.SUPPORTEDArticle 10 'Data and data governance'
41SARAHAI Act Article 10 binds providers of high-risk systems.SUPPORTEDArt. 10(1) on high-risk systems; Art. 16(a) binds providers to Section 2 requirements; Art. 10(5) names providers
43SARAHAI Act Article 10(2): training, validation and testing data sets subject to governance practices covering design choices; data collection processes and origin; preparation operations (annotation, labelling, cleaning, enrichment, aggregation); formulation of assumptions about what the data measure and represent; assessment of availability, quantity and suitability; examination for biases; identification of data gaps or shortcomings and how they will be addressed.SUPPORTEDArt. 10(2)(a)–(h)
45SARAHAI Act Article 10(3): data sets shall be 'relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose'; 10(4): take into account the specific geographical, contextual, behavioural or functional setting.SUPPORTEDArt. 10(3), verbatim; 10(4) 'geographical, contextual, behavioural or functional setting'

EU-2025-07-Annex-11-Computerised-Systems-Draft.pdf

LineSpeakerClaimVerdictEvidence
6SARAHThe Annex 11 draft's first paragraph says the GMP and GDP Inspectors Working Group and the PIC/S Committee jointly recommended the revision.SUPPORTED'Reasons for changes: The GMP/GDP Inspectors Working Group and the PIC/S Committee jointly recommended…' (p.1)
6SARAHThe Annex 11 revision grows from about five pages to nineteen, in seventeen chapters: quality system, risk management, supplier management, qualification and validation, handling of data, identity and access, audit trails, electronic signatures, periodic review, security, backup, archiving.PARTIALFooter 'Page 19 of 19'; document map 1 Scope … 17 Archiving; the current five-page Annex 11 is not in the corpus
12SARAHAnnex 11 revision principle 2.4 says data integrity is 'as defined by the ALCOA+ principles'.SUPPORTED2.4 Data integrity, verbatim
18SARAHAnnex 11 draft 4.5: QRM should assess the criticality of data to product quality, patient safety and data integrity, the vulnerability of data to deliberate or indeliberate alteration, deletion or loss, and the likelihood of detection.SUPPORTED4.5 Data integrity (p.3), near-verbatim
20SARAHAnnex 11 draft chapter 10 Handling of data has four clauses: 10.1 plausibility checks on manually entered critical data with user alerts; 10.2 transfers of critical data between systems (instrument to LIMS) on validated interfaces rather than manual transcription; 10.3 migrations on a validated process considering sending and receiving constraints; 10.4 critical data encrypted where applicable.SUPPORTED10.1 Input verification; 10.2 Data transfer ('laboratory instrument to a LIMS'); 10.3 Data migration; 10.4 Encryption — exactly four clauses
22SARAHAnnex 11 draft 11.1: all users should have unique and personal accounts; shared accounts, except read-only ones, 'constitute a violation of data integrity'; chapter 11 also covers continuous management of access, segregation of duties, least privilege, an access log and recurrent reviews.SUPPORTED11.1 (p.8), verbatim; 11.2 continuous management, 11.9 access log, 11.10 segregation of duties and least privilege, 11.11 recurrent reviews
24SARAHAnnex 11 draft chapter 12: systems where users can create, modify or delete data log every manual interaction automatically; who, what (old and new value), when and why at the time of the event; the trail is enabled and locked with any settings change itself logged; reviews per procedure by someone not involved, risk-targeted, before batch release unless a later check is justified.SUPPORTED12.1–12.8: automatic logging of manual interactions; who/what (old and new value)/when/why 'recorded at the time of events'; 12.3 'enabled and locked at all times', with deactivation by a non-GMP administrator allowed if it creates an entry; 12.8 review 'prior to batch release, unless the risk of a later detection … can be justified'
24SARAHAnnex 11 draft 12.9: a complete electronic copy must be obtainable and 'flat and locked files are not acceptable'.SUPPORTED12.9 Electronic copy, verbatim
24SARAHAnnex 11 draft chapters 16 and 17: backups at risk-based intervals, physically and logically separated, restore-tested; archived data read-only with integrity verified by checksum when moved.SUPPORTED16.2 'suitable intervals … retention determined through a risk-based approach'; 16.3/16.4 physical and logical separation; 16.6 restore test; 17.1 read-only; 17.2 'e.g. by means of a checksum'
25HOSTAnnex 11 applies to computerised systems used in GMP activities.SUPPORTED1. Scope: 'all types of computerised systems used in the manufacturing of medicinal products and active substances'
53SARAHAnnex 11 draft 6.2 says system requirements may include process maps and data flow diagrams.SUPPORTED6.2: 'Where relevant, requirements should include process maps and data flow diagrams, and use cases may be applied' ('should', stronger than the script's 'may')
70SARAHAnnex 11 draft 10.1 wants manual inputs plausibility-checked.SUPPORTED10.1 Input verification
78SARAHAnnex 11 draft 11.1 gives a unique account and 11.10 least privilege; chapter 12 audit trails by analogy.SUPPORTED11.1 unique accounts; least privilege is 11.10 Guiding principles; chapter 12 Audit Trails

EU-2025-07-Annex-22-AI-Draft.pdf

LineSpeakerClaimVerdictEvidence
8SARAHAnnex 22 describes itself as additional guidance to Annex 11 for systems with AI embedded.SUPPORTEDScope lines 7-8
16SARAHAnnex 22's test-data section requires representative of the full sample space, stratified, sufficient size, verified labels, pre-specified pre-processing, documented exclusions, and independence.SUPPORTED5.1–5.5; independence is section 6 'Test Data Independency' (the script says 'and independent', not that it sits in section 5)
50HOSTAnnex 22 5.4 wants pre-processing pre-specified with a rationale that it represents intended-use conditions.SUPPORTED5.4: 'pre-specified and a rationale should be provided, that it represents intended use conditions'
51SARAHAnnex 22 5.6: generating test data or labels by generative AI 'is not recommended and any use hereof should be fully justified'.SUPPORTED5.6, verbatim (generative AI given as 'e.g.')
65SARAHAnnex 22 3.1 and 3.2 require all common and rare variations, with subgroups by site, equipment, material and product.SUPPORTED3.1 'all common and rare variations'; subgroups by site, equipment, material and product are 3.2 (where applicable)
65SARAHAnnex 22 5.5 wants exclusions documented and justified.SUPPORTED5.5: 'Any cleaning or exclusion of test data should be documented and fully justified'
74SARAHAnnex 22 5.3: labels verified to 'a very high degree of correctness' by independent experts, validated equipment or laboratory tests; 6.4: physical vials for the final test never used in training.SUPPORTED5.3, verbatim; 6.4 'objects used for the final test … not previously been used to train or validate the model, unless features are independent' ('vials' is the script's example)

EU-2025-07-Chapter-4-Documentation-Draft.pdf

LineSpeakerClaimVerdictEvidence
10SARAHThe revised Chapter 4 is where the EU GMP guide first says the words artificial intelligence.SUPPORTEDChapter 4 draft 4.23, 4.24, 4.25 and the note after 4.36 name 'artificial intelligence'; whether it is the first place in the EU GMP guide is not checkable from the draft alone
12SARAHThe revised Chapter 4 adds a tenth attribute, traceable, and writes ALCOA++.SUPPORTEDTable 1 under 4.63 lists ten attributes with Traceable tenth; glossary 'ALCOA++'; written 'ALCOA ++' at 4.63
26SARAHChapter 4 draft 4.24: accountability for the integrity of records or raw data produced or processed with artificial intelligence rests with the regulated user.SUPPORTED4.24 (pp.3-4), verbatim ('documents, records or (raw) data')
26SARAHChapter 4 draft 4.25: support by automatic means (validation scripts or artificial intelligence) should be included in the PQS whether on premise or hosted.SUPPORTED4.25 (p.4), verbatim
26SARAHChapter 4 draft 4.27: for electronic records the regulated user should define which data are raw data, and at least all data on which quality decisions are based should be defined as raw data.SUPPORTED4.27(iv) Record (p.4), verbatim
53SARAHChapter 4 draft says for derived data 'the traceability which allows reconstruction of all data processing activities should be maintained'.SUPPORTED4.12(iii), verbatim
57SARAHChapter 4 draft wants data ownership defined throughout the lifecycle.SUPPORTED4.15: 'Data governance systems should address data ownership throughout the entire lifecycle'; 4.80
76SARAHChapter 4 draft 4.51 and 4.54: documents uniquely identifiable with a defined effective date; issuance, revision, superseding and withdrawal controlled with revision histories.SUPPORTED4.51 and 4.54 (p.9), verbatim
80SARAHChapter 4 draft 4.24 keeps accountability for AI-processed records with the regulated user.SUPPORTED4.24, verbatim

EU-2025-07-GMP-Consultation-Page-Ch4-Annex11-Annex22.pdf

LineSpeakerClaimVerdictEvidence
4SARAHThe Annex 11 revision was consulted July to October 2025 in the same package as Annex 22 and the revised Chapter 4.NOT-CHECKABLENeither draft carries a date; landscape row (web-verified)

EU-2026-07-Digital-Omnibus-AI-Reg-2026-1744.pdf

LineSpeakerClaimVerdictEvidence
41SARAHAI Act high-risk obligations are phased in through 2027 and 2028 (hedged 'as I understand it').NOT-CHECKABLELandscape §3.3 (Digital Omnibus); the AI Act text on disk still says 'apply from 2 August 2026'

FDA-2018-12-Data-Integrity-CGMP-QA.pdf

LineSpeakerClaimVerdictEvidence
4SARAHFDA's Data Integrity and Compliance with Drug CGMP is a question-and-answer guidance, final December 2018, from CDER with CBER and CVM.SUPPORTEDTitle page: CDER/CBER/CVM, December 2018; Section III 'Questions and Answers'
12SARAHFDA's Q&A question 1 says data integrity refers to the completeness, consistency and accuracy of data, and such data should be attributable, legible, contemporaneously recorded, original or a true copy, and accurate.SUPPORTEDQ1(a) p.4, verbatim
29SARAHFDA Q&A Q1: metadata is 'the contextual information required to understand data'; the number 23 is meaningless without the unit mg; includes timestamp, user, instrument, material status, audit trails; relationships between data and metadata preserved in a secure and traceable manner.SUPPORTEDQ1(b) p.4, verbatim
29SARAHFDA Q&A defines an audit trail as a secure, computer-generated, time-stamped record that allows the course of events to be reconstructed.SUPPORTEDQ1(c) p.4: 'a secure, computer-generated, time-stamped electronic record that allows for reconstruction of the course of events'
29SARAHFDA Q&A: a static record is fixed (paper or image); a dynamic record lets the user interact with the content, example reprocessing a chromatogram.SUPPORTEDQ1(d) p.5
31SARAHFDA Q&A Q2: a result may be invalidated and excluded from the batch decision only with 'a valid, documented, scientifically sound justification', with the original data retained alongside the investigation.SUPPORTEDQ2 p.6, verbatim
31SARAHFDA Q&A Q3: each CGMP workflow on a computer system is validated for intended use; qualifying the platform does not show the workflow computes correctly; example an execution system whose master record might have the wrong calculation.SUPPORTEDQ3 p.6 (MES / MPCR example)
33SARAHFDA Q&A Q4 and Q5: access restricted to authorised people, technically where possible, administrators independent of record owners, a list of authorised persons; a shared login means no unique individual can be identified so the system would not conform.SUPPORTEDQ4 p.7; Q5 p.7: 'a unique individual cannot be identified through the login and the system would not conform'
33SARAHFDA Q&A Q8: audit trails reviewed at the frequency the regulation sets for the record, otherwise at a risk-based frequency weighing data criticality, controls and impact on product quality.SUPPORTEDQ8 p.8
35SARAHFDA Q&A Q12: 'when generated to satisfy a CGMP requirement, all data become a CGMP record'; saved at the time of performance; processes designed so required data cannot be modified without a record of the modification.SUPPORTEDQ12 p.10, verbatim
35SARAHA footnote to FDA Q&A Q12 says FDA routinely requests records not intended to satisfy a requirement but containing CGMP information.SUPPORTEDFootnote 14 to Q12: 'FDA routinely requests and reviews records not intended to satisfy a CGMP requirement but which nonetheless contain CGMP information'
59SARAHFDA Q&A lists the instrument identifier among metadata required to reconstruct the activity.SUPPORTEDQ1(b): 'the instrument ID used to acquire the data'
74SARAHFDA Q&A Q4 names automated visual inspection records among records whose alteration should be restricted to authorised personnel.SUPPORTEDQ4 p.7: 'Other examples of records for which control should be restricted to authorized personnel include automated visual inspection records'

FDA-2025-01-AI-Regulatory-Decision-Making-Draft.pdf

LineSpeakerClaimVerdictEvidence
14SARAHFDA 2025 step 4.a: development data should be 'fit for use', defined as relevant and reliable; relevant illustrated as key data elements and sufficient data representative of the manufacturing process or operation; reliable as accurate, complete, and traceable; step 4.b says the same of test data.SUPPORTED4.a.ii lines 336-339; 4.b line 412 'Like development data, these data should be fit for use'
39SARAHFDA 2025 step 4.a asks to describe development datasets and their split into training and tuning; collection, processing, annotation, storage and control; rationale; how labels were established; how the data are fit for the COU.SUPPORTED4.a.ii lines 349-360
39SARAHFDA 2025 step 4.b has a clause about what the draft calls data drift: describe the applicability of the test data to the COU, because a model may not perform as well if development data differ from data in the deployed environment.SUPPORTED4.b lines 436-440: '…This phenomenon is sometimes referred to as data drift.'

PICS-2021-07-PI-041-Data-Integrity.pdf

LineSpeakerClaimVerdictEvidence
10SARAHPIC/S PI 041 on data management and integrity is dated 1 July 2021 and written for inspectors.SUPPORTEDCover 'PI 041-1, 1 July 2021'; 3.1.1 'guidance for Inspectorates'
12SARAHPI 041 adds complete, consistent, enduring and available and calls it ALCOA+.SUPPORTED7.4: 'Complete, Consistent, Enduring and Available (ALCOA+)'
37SARAHPI 041's data lifecycle runs generated, processed, reported and checked, used for decision-making, stored, discarded; data cross boundaries (paper to computer, production to QC to QA, contract giver to acceptor).SUPPORTED5.1.2, verbatim
37SARAHPI 041 grades data by criticality (which decision the data influence) and risk (opportunity for alteration or deletion and likelihood of detection).SUPPORTED5.4.1 'Which decision does the data influence?'; 5.5.2 vulnerability and likelihood of detection
37SARAHPI 041's data-risk factors include multi-stage processes, transfer between systems, complex processing, and 'biological production processes or analytical tests may exhibit a higher degree of variability compared to small molecule chemistry'.SUPPORTED5.5.4, verbatim
38HOSTPI 041 says an organisation that believes there is no risk of data-integrity failure is unlikely to have assessed its data lifecycle.SUPPORTED5.5.6, verbatim
57SARAHPI 041 says an indicator of data governance maturity is an organisational understanding and acceptance of residual risk.SUPPORTED5.5.6, verbatim
59SARAHPI 041's ALCOA table: attributable means it should be possible to identify the individual or computerised system that performed the task and when.SUPPORTED7.5 table, Attributable, verbatim
61SARAHPI 041: complete means all information critical to recreating an event, including metadata for electronic data.SUPPORTED7.5 table, Complete
63SARAHPI 041: accurate means a truthful representation of facts, achieved through qualified and calibrated equipment and validated systems, procedures and data review, deviation management, and trained people.SUPPORTED7.5 table, Accurate: 'truthful representation of facts … This can be comprised of: equipment related factors such as qualification, calibration, maintenance and computer validation; policies and procedures …; deviation management …; trained and qualified personnel'
82SARAHPI 041 suggests a data custodian as a role an organisation might create.SUPPORTED6.6.6: 'new roles … such as a data custodian might be considered'
83HOSTPI 041 spends a whole section on culture.SUPPORTED6.3 'Quality culture' (6.3.1–6.3.3)

Host opinions (listed, not verified)