AI in Biopharma Manufacturing: Are We There Yet?
This is the script the AI voices read, so it matches the audio word for word; times are from the render. Researched, scripted and voiced by AI systems under Jack Prior's direction. Sam and Sarah are AI characters; nothing in the episode is Jack speaking, and none of it is a statement of his views or his employer's. Generative AI can be confidently wrong — check the sources. Corrections: jack@jackprior.ai.
00:00 Narrator This is Are We There Yet — a podcast on the evolution of AI in biopharma manufacturing, directed by Jack Prior. A word about how it's made. This episode was researched, scripted and voiced by AI systems. Jack sets the questions and the frames; the AI reads the documents and writes the conversation you're about to hear, between two AI characters: Sam, who plays a manufacturing-science practitioner, and Sarah, who has read the documents. Nothing you hear is Jack speaking, and none of it is a statement of his views or his employer's. Like any generative AI output, it can be wrong — confidently wrong, or missing a nuance — which are exactly the risks this industry is working to mitigate, and exactly what this season is about. Check the sources before you rely on anything. Corrections are welcome at jackprior dot A I. Now, the episode.
00:53 Sam A biologics plant. The data science team pulls three years of bioreactor data from the historian to train a soft sensor for titre. Forty thousand rows, a model by Friday. Then the validation lead asks four questions. Which probe was this pH signal from in year two, before the skid was rebuilt? Why are there no values between two and four one morning in March? Are these raw readings, or what the historian kept after discarding every change inside its deadband? If I delete this file, can you regenerate it? Not sure, an outage nobody logged, the second one, and no. The training set has just become a story about data, and no framework this season accepts a story.
01:37 Sarah That's tonight's question, Sam. Every document we've read assumes the data under the model are trustworthy. Tonight we read what they actually require, and where that leaves a plant.
01:47 Sam And where it leaves the basement. Let's read.
01:50 Sam Sarah, the question in one sentence.
01:52 Sarah Every framework presumes trustworthy data: what do they actually require, and where does that leave most plants? Two anchors tonight, and both come from the data-integrity rulebook rather than the AI one. The revision of EU GMP Annex Eleven on computerised systems, a draft consulted from July to October twenty twenty-five, in the same package as Annex Twenty-two and the revised Chapter Four on documentation. And FDA's Data Integrity and Compliance with Drug CGMP, a question-and-answer guidance, final in December twenty eighteen, from the Center for Drug Evaluation and Research with the biologics and veterinary centres.
02:32 Sam Provenance and maturity.
02:34 Sarah Annex Eleven is the law-and-regulation layer, an annex to EudraLex Volume Four. The draft's own first paragraph says the GMP and GDP Inspectors Working Group and the PIC/S Committee, the Pharmaceutical Inspection Co-operation Scheme, jointly recommended the revision. It grows from about five pages to nineteen, in seventeen chapters: quality system, risk management, supplier management, qualification and validation, handling of data, identity and access, audit trails, electronic signatures, periodic review, security, backup, archiving. As I understand it, the final text is expected around the end of this year and implementation around early twenty twenty-seven; treat both as estimates.
03:24 Sam And Annex Twenty-two hangs off it.
03:26 Sarah It describes itself as additional guidance to Annex Eleven for systems with AI embedded, so the two are read together. FDA's Q and A is regulator guidance, final, seven years old, and it is what an FDA investigator has in mind whenever the word data comes up.
03:42 Sam Voices.
03:43 Sarah Five. PIC/S guidance PI zero-four-one on data management and integrity, dated the first of July twenty twenty-one, written for inspectors. The revised Chapter Four, same consultation, which is where the EU GMP guide first says the word artificial intelligence. FDA's January twenty twenty-five AI draft, for its definition of fit for use. The EU AI Act, in force since August twenty twenty-four, for Article Ten on data governance. And EMA's reflection paper on AI across the medicinal product lifecycle, final September twenty twenty-four, for its section on data acquisition.
04:24 Sam Why now, in one breath. The soft sensor and the batch model have lived on historian data for twenty years, and the data-integrity rulebook grew up around the lab, the chromatograph and the batch record; the historian mostly slid underneath it. What changed is the agent. An assistant that reads batch records, procedures and deviation history doesn't consume a column of numbers. It consumes the plant's entire documentary memory, versions and all, and the moment you put a language model on top of that, every weakness in the basement is on the screen. Last time we asked what evidence you owe. Tonight we ask whether the data can carry it.
05:03 Sarah Start with vocabulary, because the corpus defines trustworthy data in three overlapping ways. The first is ALCOA. FDA's Q and A, question one, says data integrity refers to the completeness, consistency and accuracy of data, and that such data should be attributable, legible, contemporaneously recorded, original or a true copy, and accurate. PIC/S adds four attributes, complete, consistent, enduring and available, and calls it ALCOA plus. Annex Eleven's revision says in principle two point four that data integrity is, quoting, as defined by the ALCOA+ principles. And the revised Chapter Four adds a tenth, traceable, and writes ALCOA plus plus.
05:54 Sam Three texts, three spellings, one idea. Reliability.
05:57 Sarah The second definition is FDA's twenty twenty-five draft, in step four A. The data used to develop the model should be, quoting, fit for use, which the draft defines as both relevant and reliable. Relevant is illustrated as including key data elements and sufficient data that are representative of the manufacturing process or operation. Reliable is defined as accurate, complete, and traceable. And the draft says the same of test data in step four B: like development data, these data should be fit for use.
06:30 Sam So reliable is ALCOA in three words, and relevant is the new word.
06:33 Sarah The third is Annex Twenty-two's test-data section, which we read last time: representative of the full sample space, stratified, sufficient in size, labels verified, pre-processing pre-specified, exclusions documented, and independent. That is the relevance idea written as a recipe for one dataset.
06:53 Sam Here's what I want listeners to hold. Reliability is the old agenda; quality has been enforcing ALCOA for a decade. Relevance is the new one: does this dataset represent the process the model will meet? No audit trail answers that. And in the kind of plant we're describing, relevance is where the data fail. The records are attributable, contemporaneous and utterly unrepresentative, because they came from the campaigns that happened to be convenient.
07:20 Sarah Annex Eleven is where reliability becomes explicit GMP for any system, and I'll read it for what it does to a training pipeline. Four point five is the data clause of risk management: quality risk management should be used to assess the criticality of data to product quality, patient safety and data integrity, the vulnerability of data to deliberate or indeliberate alteration, deletion or loss, and the likelihood of detecting it.
07:49 Sam Chapter ten.
07:51 Sarah Handling of data, four clauses. Ten point one: where critical data are entered manually, the system should check plausibility, for instance against expected ranges, and alert the user. Ten point two: transfers of critical data between systems, its example is instrument to LIMS, the laboratory information management system, should run on validated interfaces rather than manual transcription. Ten point three: migrations on a validated process, considering the constraints on the sending and receiving side. Ten point four: critical data encrypted where applicable.
08:30 Sam Ten point two is my whole pipeline. Historian to a data lake to a notebook is three transfers, and two of them are usually a script somebody wrote on a Thursday.
08:38 Sarah Chapter eleven, identity and access. Eleven point one: all users should have unique and personal accounts, and shared accounts, except read-only ones, quoting, constitute a violation of data integrity. Then continuous management of access, segregation of duties, least privilege, an access log, and recurrent reviews.
09:01 Sam And audit trails.
09:02 Sarah Chapter twelve. Systems where users can create, modify or delete data should log every manual interaction automatically; the trail captures who, what including the old and new value, when, and why, recorded at the time of the event and not at the end of a process; it is enabled and locked, and any change to its settings is itself logged; reviews follow a procedure, are done by someone not involved in the activity, are targeted by risk, and happen before batch release unless a later check is justified. Twelve point nine: a complete electronic copy must be obtainable, and, quoting, flat and locked files are not acceptable. Chapters sixteen and seventeen: backups at risk-based intervals, physically and logically separated, restore-tested; archived data read-only, their integrity verified by checksum when they move.
10:03 Sam Now the consequence for us. If the training set is a GMP record, and under FDA's question twelve it becomes one the moment it serves a GMP decision, then the pipeline that built it is a computerised system used in GMP activities, and this annex applies to it. Unique accounts on the data platform. An audit trail on the feature table. Backups of the training set. Most data science environments are built for speed, with a shared service account and a notebook, and the annex says the shared account is a data-integrity violation on its own.
10:42 Sarah And the revised Chapter Four names AI directly. Four point twenty-four: accountability for the integrity of records or raw data produced or processed with artificial intelligence rests with the regulated user. Four point twenty-five: support by any automatic means, its examples are validation scripts or artificial intelligence, should be included in the pharmaceutical quality system whether it runs on premise or as a hosted service. And four point twenty-seven says that for electronic records the regulated user should define which data are raw data, and at least all data on which quality decisions are based should be defined as raw data.
11:23 Sam So the site has to write down, in advance, that the historian tags feeding the soft sensor are raw data. Which is a decision, with an owner.
11:32 Sam FDA's Q and A, in one breath, because the audience knows ALCOA and I don't want to relitigate it. Give me the questions that matter for a model.
11:40 Sarah Question one defines the terms. Data integrity we have. Metadata is, quoting, the contextual information required to understand data; the number twenty-three is meaningless without the unit milligrams; it includes the timestamp, the user, the instrument, material status and audit trails; and the relationships between data and metadata should be preserved in a secure and traceable manner. An audit trail is a secure, computer-generated, time-stamped record that allows the course of events to be reconstructed. Static versus dynamic: a static record is fixed, a paper or an image; a dynamic record lets the user interact with the content, and the example is reprocessing a chromatogram.
12:30 Sam A historian tag with its compression settings is a dynamic record. Re-query it with different settings and you get different numbers.
12:37 Sarah Question two: when may a result be invalidated and excluded from the batch decision? Only with, quoting, a valid, documented, scientifically sound justification, and even then the original data stay in the record alongside the investigation. Question three: each CGMP workflow on a computer system is validated for its intended use; qualifying the platform does not show the workflow computes correctly, and the example is an execution system whose master record might still have the wrong calculation.
13:11 Sam Access.
13:12 Sarah Questions four and five: access restricted to authorised people, technically where possible, with administrators independent of the record owners and a list of who is authorised; and a shared login means no unique individual can be identified, so the system would not conform. Question eight: audit trails reviewed at the frequency the regulation sets for the record, otherwise at a risk-based frequency that weighs data criticality, the controls in place and the impact on product quality.
13:45 Sam And twelve.
13:46 Sarah Question twelve asks when electronic data become a CGMP record. The answer, quoting: when generated to satisfy a CGMP requirement, all data become a CGMP record. They must be saved at the time of performance, and processes designed so that required data cannot be modified without a record of the modification. And a footnote adds that FDA routinely requests records that were never intended to satisfy a requirement but nonetheless contain CGMP information.
14:19 Sam That footnote is the quiet one. The training set on the shared drive wasn't intended to satisfy anything. It contains process data. They can ask for it.
14:28 Sarah PI zero-four-one is the inspector's version of the same expectations, written for on-site inspections. Its data lifecycle runs from generated, processed, reported and checked, through used for decision-making, to stored and finally discarded, and it notes that data cross boundaries: paper to computer, production to quality control to quality assurance, contract giver to acceptor. It grades data two ways: criticality, which decision the data influence, and risk, the opportunity for alteration or deletion and the likelihood of detection. Its list of things that raise data risk includes multi-stage processes, transfer between systems, complex processing, and, quoting, biological production processes or analytical tests may exhibit a higher degree of variability compared to small molecule chemistry.
15:20 Sam That's a bioreactor, named by an inspectors' guide as a data-risk factor before any model exists. And one line I'll keep: an organisation that believes there is no risk of data-integrity failure is unlikely to have assessed its data lifecycle. So far, then. Three definitions: ALCOA and its pluses, which is reliability; FDA's fit for use, which adds relevance; Annex Twenty-two's test-data recipe. Annex Eleven makes the pipeline a GMP system: validated transfers, plausibility checks, unique accounts, audit trails, backup, archive. FDA says the data became a record the moment they served a GMP decision, and excluding any of them needs a written justification. Now the AI texts.
16:13 Sarah Four documents add relevance and provenance. FDA's draft first, step four A: describe the development datasets and how they were split into training and tuning; how they were collected, processed, annotated, stored and controlled; the rationale for choosing them; how labels were established; and how the data are fit for the context of use. For test data the same list, plus a clause about what the draft calls data drift: describe the applicability of the test data to the context of use, because a model may not perform as well if the development data differ from the data it meets in the deployed environment.
16:56 Sam Applicability. Third time this season that word has turned out to be the whole game.
17:01 Sarah The AI Act next, with the caveat from episode three: Article Ten binds providers of high-risk systems, which most pharma manufacturing AI is not, and, as I understand it, those obligations are now phased in through twenty twenty-seven and twenty twenty-eight. But it is the clearest statement anywhere of what data governance for a model means.
17:25 Sam Paragraph two.
17:26 Sarah Training, validation and testing data sets shall be subject to governance practices covering the design choices; data collection processes and the origin of the data; preparation operations such as annotation, labelling, cleaning, enrichment and aggregation; the formulation of assumptions, in particular about what the data are supposed to measure and represent; an assessment of the availability, quantity and suitability of the data sets; examination for biases; and the identification of data gaps or shortcomings and how they will be addressed.
17:58 Sam And paragraph three.
17:59 Sarah The data sets shall be, quoting, relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose. Paragraph four: they take into account the specific geographical, contextual, behavioural or functional setting the system is intended for.
18:18 Sam Read that as a plant. Origin of data: which historian, which tags. Assumptions about what the data measure: that the pH probe was reading pH. Setting: this line, this scale, this product. Gaps and how you'll address them. It's a data-readiness assessment, sitting in a regulation that mostly doesn't apply to us, and it's better than anything in the ones that do.
18:41 Sarah EMA's reflection paper, section two point five point one. Models are intrinsically data-driven and so vulnerable to bias; acquire a balanced and sufficiently large training set in relation to the intended context of use; and the sources of data, the acquisition process and any processing, it lists cleaning, transformation, imputation, annotation, normalisation and augmentation, should be, quoting, documented in a detailed and fully traceable manner in line with GxP requirements. Exploratory analysis should record relevance and representativeness, the interpolation and extrapolation assumptions made, and class imbalances. Augmentation may be applied to expand the training set, and any limitations that remain are to be presented in the model documentation.
19:33 Sam And on the split.
19:34 Sarah Two point five point two: an early train-test split before any normalisation is strongly encouraged, because the risk of data leakage, quoting, often unintentional or even unconscious, cannot be completely excluded.
19:47 Sam Imputation. That's the historian gap filled by interpolation. EMA wants it documented and traceable; Annex Twenty-two five point four wants pre-processing pre-specified, with a rationale that it represents intended-use conditions. And on synthetic data?
20:04 Sarah A divergence worth stating. EMA says augmentation and synthetic data may in some cases be useful for expanding the training set. Annex Twenty-two five point six says generating test data or labels by generative AI, quoting, is not recommended and any use hereof should be fully justified. They speak of different sets, training and test, so a site can hold both positions: augment the training set if you document it; never let generated data into the test set.
20:35 Sam Now the gap. Sarah, does any of these documents say how you get from a historian, a batch record system and a LIMS to a dataset that meets any of that?
20:42 Sarah No. The nearest is one sentence in EMA's section on model deployment: the data acquisition hardware, software and data transformation pipeline at inference should be in line with pre-defined specifications. Annex Eleven six point two says system requirements may include process maps and data flow diagrams. Chapter Four's lifecycle asks that for derived data, quoting, the traceability which allows reconstruction of all data processing activities should be maintained. Those are the three sentences, and none of them says what the pipeline has to do.
21:21 Sam Then let me say what it has to do, from the practitioner's chair. Contextualisation: which batch, which phase, which unit operation, which equipment train this tag was on at this timestamp. Time alignment: a lab result at nine in the morning belongs to a sample drawn at seven, from a reactor whose online signals log every ten seconds and whose feeds log only when they change. Units: the same tag in grams per litre on one skid and milligrams per millilitre on another. And the metadata that makes a value relevant at all: probe identity, calibration state, controller mode. None of that is in the documents, and all of it decides whether the fit-for-use sentence is true.
22:05 Sarah The documents assume it. Annex Twenty-two's input sample space, FDA's representative of the manufacturing process, the AI Act's what the data are supposed to measure and represent: each presumes someone has already done that work.
22:20 Sam Here's my view, labelled as mine. Most AI readiness gaps are data readiness gaps. When a use case stalls, it's rarely the algorithm; it's that nobody can say which batches these rows came from, or whether the process the model learned is the process it will meet. A maturity scale for data readiness, where your data sit on the road from raw historian to contextualised, aligned, traceable and relevant, would be worth more to most sites than another AI policy. It's the show's question, asked of the data first. Are we there yet? Usually not, and you can't fix it with the model.
22:57 Sarah Two things in the documents support that. PI zero-four-one says an indicator of data governance maturity is an organisational understanding and acceptance of residual risk. And Chapter Four wants data ownership defined throughout the lifecycle. Ownership, and honesty about the gaps, are both on the page.
23:16 Sam Restate before the examples. The AI texts add relevance and provenance: FDA wants development and test data both fit for use, with their applicability described; the AI Act, where it applies, wants origin, assumptions, suitability and gaps written down; EMA wants acquisition and every processing step traceable; Annex Twenty-two wants nothing generated in the test set. Nobody says how to build the pipeline. Now the soft sensor's training set, one attribute at a time.
23:49 Sarah Take the file from the cold open and interrogate it. Attributable and contemporaneous first. PI zero-four-one's table asks that it be possible to identify the individual or computerised system that performed the task and when. For a historian value the who is an instrument: which probe, which transmitter, which controller mode, and its calibration state that day. If the tag name survived a skid rebuild but the probe did not, the column is two instruments wearing one name. FDA's metadata answer lists the instrument identifier as metadata required to reconstruct the activity; if it isn't in the dataset, the dataset isn't complete.
24:34 Sam Complete.
24:35 Sarah PI zero-four-one: all information that would be critical to recreating an event, and for electronic data that includes the metadata. So the outage matters twice. The missing hours are an incompleteness. And if the historian's compression discarded every value inside a deadband, what was stored is already a derived record, which Chapter Four says must carry the traceability to reconstruct the processing. FDA's dynamic-record idea applies too: a record you can re-query with different settings and get different numbers from is dynamic, and a flat export of it is not the original.
25:11 Sam Accurate.
25:12 Sarah A truthful representation of facts, achieved, the table says, through qualified and calibrated equipment and validated systems, procedures and data review, deviation management, and trained people. For a soft sensor the label is a laboratory assay, so accuracy also means the assay's own method performance, which is the reference-method summary FDA asked for last time.
25:36 Sam Relevant. This is the one.
25:38 Sarah FDA: representative of the manufacturing process or operation. Annex Twenty-two three point one and three point two: all common and rare variations, with subgroups by site, equipment, material and product. The AI Act: the specific functional setting. So: which campaigns are in the file, which products, which changeovers, which scales, which seasons; whether the process the model learned is the process it will see; and whether the rare variations, a contaminated run, a feed pump fault, are in there at all, or were the batches nobody exported. Leaving those batches out is an exclusion decision, and FDA's question two and Annex Twenty-two five point five both want it documented and justified.
26:28 Sam Traceable.
26:30 Sarah Chapter Four's tenth attribute and FDA's third word for reliable: can the dataset be regenerated from raw records. If not, the training set is an original with no lineage, and Annex Eleven's chapters on backup and archiving apply to it as a record in its own right.
26:47 Sam So the soft sensor's training set, honestly assessed, is usually attributable in name only, contemporaneous but compressed, complete except where it isn't, accurate to the assay, relevant to the campaigns somebody chose, and traceable to a script that no longer runs. That's the basement. The model on top was never the problem.
27:07 Sam The batch model.
27:08 Sarah Multivariate batch monitoring adds two data problems and one decision. Batch context: every row has to carry its batch identity and its phase, and phase boundaries in a historian are often a manual entry or an inferred trigger, which is exactly the manual input Annex Eleven ten point one wants plausibility-checked. Phase alignment: stretching batches of unequal length onto a common axis is a pre-processing step, pre-specified with a rationale under Annex Twenty-two five point four.
27:42 Sam And the golden batches.
27:44 Sarah The reference batches the model is built from. Every batch left out of that set is an exclusion, and the two documents agree: FDA's question two, only with a valid, documented, scientifically sound justification, with the excluded data retained; Annex Twenty-two five point five, documented and fully justified.
28:05 Sam In practice the golden batches are the ones the process engineer liked. That's a scientific judgement, and it can be a good one. It just has to be written down at the time, not reconstructed for the inspector. The camera.
28:19 Sarah The vision system's data are images with labels, so the two clauses that bite are Annex Twenty-two five point three, labels verified to, quoting, a very high degree of correctness, by independent experts, validated equipment or laboratory tests, and six point four, the physical vials used for the final test never used in training. Under ALCOA, attributable means which camera, which lighting configuration, which line. And FDA's question four names automated visual inspection records among the records whose alteration should be restricted to authorised personnel, so the image store is a controlled record, not a folder.
29:00 Sam And the agent. Slowly, because for the agent, data are documents.
29:04 Sarah For the agent the input is batch records, procedures and deviation history, and no document in the corpus addresses that directly. So apply the nearest ones. Fit for use, relevant: is the procedure the agent reads the current version? Chapter Four four point fifty-one and four point fifty-four: documents uniquely identifiable, with a defined effective date, and their issuance, revision, superseding and withdrawal controlled with revision histories. A shared drive holding three versions of one procedure fails that before any model runs.
29:43 Sam Reliable, and access.
29:44 Sarah Reliable: are the batch records it reads originals or true copies, and are the deviations it learns from closed, signed and complete. Access: Annex Eleven eleven point one, a unique account for the agent itself, and eleven point ten, least privilege, so that what it may read is a decision someone made and can revoke. And audit trail: Annex Eleven chapter twelve, by analogy, a record of what it read to reach a conclusion, which record, which version, when.
30:21 Sam That last one is the real ask. For a soft sensor the audit trail is on the data. For an agent the audit trail has to be on the reading. If the drafted investigation cites a procedure, I want to know which version it retrieved, logged the way a manual user's interactions would be. Nothing says that. Annex Eleven's chapters ten to twelve by analogy is the best answer anyone has today, and it's a decent one, because the discipline already exists for people.
30:47 Sarah And Chapter Four four point twenty-four keeps the accountability where it was. For records produced or processed with artificial intelligence, it rests with the regulated user.
30:57 Sam So what for biomanufacturing. Opinions labelled. If you do one thing: write the data-readiness assessment before the model. A documented description, owned by a process expert, of where the data come from, what they represent, which campaigns and equipment and products, what's missing and why, and how it can be regenerated. That one document is Annex Twenty-two's input sample space, FDA's fit-for-use rationale and the AI Act's Article Ten list at once. It's the artefact a regulator reads first, and it's the one almost nobody has.
31:38 Sarah The documents support the ownership. Annex Twenty-two puts a process subject matter expert in charge of the intended-use description; Chapter Four asks for data ownership throughout the lifecycle; PI zero-four-one suggests a data custodian as a role an organisation might create.
31:57 Sam Second, who has to learn what. Quality already owns ALCOA. Data science has to learn it, and that means unique accounts, audit trails and a controlled test-set repository in the environment they actually work in, which most of them will feel as friction. Quality, in turn, has to learn relevance. A reviewer who only checks attributability will pass a beautifully documented, unrepresentative training set. Both moves are cultural, and PI zero-four-one spends a whole section on culture for a reason.
32:27 Sarah Where the teams feel it. Manufacturing science and process science own contextualisation, the campaign and equipment map, and the exclusion decisions. Automation and IT own the pipeline as a GMP system: validated interfaces, access, backup. Quality owns the raw-data definition under Chapter Four, test-set custody, and audit-trail review. Regulatory CMC decides how much of this goes into a filing as the fit-for-use description.
32:57 Sam Third, the basement and the kitchen. The model is the kitchen: visible, exciting, where the budget goes. The pipeline is the basement, and the house stands on it. My view: spend on the basement first, because every model you build afterwards uses it, and it is the agent, not the soft sensor, that will expose whether it's dry. A soft sensor can hide a bad historian behind an error bar. An assistant that cites the wrong version of a procedure is a data problem you can see.
33:26 Sam Three things. Sarah.
33:28 Sarah First, trustworthiness has two halves. Reliability is ALCOA and its pluses, in FDA's Q and A, PI zero-four-one, Annex Eleven and Chapter Four. Relevance is FDA's fit for use, Annex Twenty-two's sample space and the AI Act's Article Ten: do the data represent the process the model will meet. Reliability is the old agenda; relevance is the new one, and the one plants fail.
33:58 Sam Second, mine: the pipeline is a GMP system. The moment historian data serve a GMP decision they're a record, per FDA's question twelve, and the Annex Eleven revision then reaches everything that touches them: validated transfers, plausibility checks, unique accounts, audit trails, backup, archive. Exclusions are decisions with a written justification. And nobody tells you how to build the pipeline. That's the gap, and most AI readiness gaps live in it.
34:29 Sarah Third, for the agent, data are documents, and the corpus is silent. The best available answer is Annex Eleven's chapters on data handling, access and audit trails read by analogy, with Chapter Four's document control: current versions, an account of its own, and a log of what it read.
34:47 Sam Next time: humans. Who is accountable when the model is wrong? Every document tonight had a person in it, the process expert, the data owner, the reviewer, and last time FDA asked us to evaluate the human-AI team rather than the model alone. Next episode reads what human oversight actually means across the documents, and takes up the question we've parked since episode four: which human, with which competence. Are we there yet? The data can be made to carry the evidence. Next we ask who is standing on it.