AI in Biopharma Manufacturing: Are We There Yet?

Transcript — Evidence: What Do I Have to Show?

This is the script the AI voices read, so it matches the audio word for word; times are from the render. Researched, scripted and voiced by AI systems under Jack Prior's direction. Sam and Sarah are AI characters; nothing in the episode is Jack speaking, and none of it is a statement of his views or his employer's. Generative AI can be confidently wrong — check the sources. Corrections: jack@jackprior.ai.

00:00 Narrator This is Are We There Yet — a podcast on the evolution of AI in biopharma manufacturing, directed by Jack Prior. A word about how it's made. This episode was researched, scripted and voiced by AI systems. Jack sets the questions and the frames; the AI reads the documents and writes the conversation you're about to hear, between two AI characters: Sam, who plays a manufacturing-science practitioner, and Sarah, who has read the documents. Nothing you hear is Jack speaking, and none of it is a statement of his views or his employer's. Like any generative AI output, it can be wrong — confidently wrong, or missing a nuance — which are exactly the risks this industry is working to mitigate, and exactly what this season is about. Check the sources before you rely on anything. Corrections are welcome at jackprior dot A I. Now, the episode.

00:53 Sam A fill-finish line, a European inspection, day two. The site has replaced manual inspection of fill level with a vision system: validated, twelve months in service, defect catch rate on the slide. The inspector reads the validation report and asks three questions in a row. What was the catch rate of the manual inspection you replaced? Who on the data science team had access to the test images before the model was trained? And for the vials it rejected during testing, what in the image made it reject them? The validation lead has accuracy, precision and a confusion matrix. She has no number for the manual process, nobody has ever asked her who saw the test set, and the heat maps were something the vendor showed in a sales meeting. None of those were validation questions two years ago. All three are in a six-page draft annex today.

01:45 Sarah That's the frame tonight, Sam. What do I have to show? Not whether the model works, but what counts as evidence that it does, on paper, in front of someone who didn't build it.

01:55 Sam And how the word for that stopped being validation and started being assurance. Let's read.

What do I have to show, and why now

02:01 Sam Sarah, the question.

02:02 Sarah What does credibility look like on paper, and how has validation turned into assurance? Two anchors. FDA's January twenty twenty-five draft on AI to support regulatory decision-making, whose first three steps we read in episode two; tonight, step four, the credibility assessment plan, and step six, the report. And draft Annex Twenty-two, whose gate we read last time; tonight the sections behind the gate, three to nine, the most concrete evidence recipe in the corpus.

02:37 Sam Why now, in a sentence, because I've said the long version. The soft sensor and the batch model were validated for twenty years against a rulebook that said validate and left the rest to the site. What changed is that the rulebook started writing down what the evidence is, in the same two years the agentic assistant walked in with no idea what its evidence would even look like. The recipe arrived exactly when the hardest case did.

03:01 Sam Provenance and maturity.

03:03 Sarah FDA's draft: regulator guidance from CDER with CBER, CDRH and other centres, January twenty twenty-five, still a draft, manufacturing named in its scope. Annex Twenty-two: law-and-regulation layer, an annex to EudraLex Volume Four, drafted by EMA's inspectors working group and consulted with PIC/S, out for comment July to October twenty twenty-five, roughly thirteen hundred comments, no revised text published after the EMA workshop this summer, though EMA says it is revising the draft. As I understand it, final to the Commission around the end of this year and effective around twenty twenty-seven; treat those as estimates. Both are drafts. Both are what a reviewer or inspector reads today.

03:54 Sam Voices.

03:55 Sarah Four. FDA's Computer Software Assurance guidance for production and quality management system software, device-side, final September twenty twenty-five and reissued the third of February twenty twenty-six, where the word assurance comes from. The ten principles of Good Machine Learning Practice, from FDA, Health Canada and MHRA in twenty twenty-one, adopted internationally as IMDRF document N eighty-eight in January twenty twenty-five. FDA's November twenty twenty-three device guidance on the credibility of computational modelling, where the credibility vocabulary was written down. And BioPhorum's June twenty twenty-six AI risk guidance, industry practice, for what eight companies propose as evidence once the risk is graded.

04:46 Sam And the thread from last time. The gate said which models may be there at all. Tonight is what you owe once you're through it.

FDA: plan, execute, report, commensurate

04:53 Sam Start with FDA, because FDA's answer to how much evidence is the shortest.

04:57 Sarah Commensurate with model risk. The scope section lists what scales, the level of oversight, the stringency of the assessments and the acceptance criteria, the risk mitigation, the extent of documentation, and says all of it should be, quoting, commensurate with the AI model risk and tailored to the specific COU. Step four's one concrete illustration: performance acceptance criteria should be more stringent, and described to FDA in more detail, for high-risk models than for low-risk ones. For certain low-risk models FDA may ask for minimal information; for high-risk, everything it lists and more.

05:39 Sam So the framework tells me the dial exists and which way to turn it. It doesn't tell me the setting.

05:44 Sarah It does not, and it says why: steps one to three have worked examples and step four does not, because the appropriate activities vary with the nuances of a specific programme. What it gives instead is a structure. Step four is a credibility assessment plan in two halves. Four A describes the model and how it was developed: the model, the data used to develop it, and the training. Four B describes how the model is evaluated on test data. Step five executes the plan. Step six documents the results, and any deviations from the plan, in a credibility assessment report. Step seven decides whether the model is adequate for the context of use.

06:30 Sam And where does the report go? A plant isn't filing a submission for every model.

06:35 Sarah Two homes. A self-contained document in a submission or a meeting package, or, quoting, held and made available to FDA on request, for example during an inspection. A footnote adds that for uses outside the usual meeting routes sponsors may complete all seven steps without seeking early engagement. So the plan-then-report discipline can live in the site's quality system, with the report held for inspection.

07:01 Sam Which is my first takeaway, as the practitioner. Nobody is going to tell you how much is enough. What they will ask is whether you wrote the plan before you ran the tests, whether the report matches the plan, and whether the gap between the two is explained. The discipline is the evidence. A site with a beautiful accuracy number and no plan has a result.

Annex 22: intended use owned by a process expert

07:21 Sarah Annex Twenty-two is where the structure gets concrete. A draft on the binding track, a supplement to the computerised-systems annex, six pages, and after the gate in section one it is almost entirely about evidence. I'll group its sections by idea rather than read them in order. The first idea is that evidence starts before the model, with the intended use.

07:45 Sam Which we'd call context of use.

07:47 Sarah Same artefact, different word. Section three says the intended use, and the specific tasks the model is designed to assist or automate, should be described in detail, based on in-depth knowledge of the process the model is integrated in. The description must contain a comprehensive characterisation of the data the model will use as input, and all common and rare variations. The draft's name for that is, quoting, the input sample space. Limitations and possible erroneous or biased inputs are to be identified. A process subject matter expert is responsible for the adequacy of the description, and it is documented and approved before acceptance testing starts.

08:30 Sam A process expert. Not the data scientist, not the vendor. That's the right design choice. The data scientist knows what the data looked like; the process person knows what it can look like: the campaign on the old probe, the supplier change nobody flagged, the lighting after the line was moved. Rare variations are the ones the training set never saw. And it makes the intended-use document a process document, which means MSAT has to write it.

08:55 Sarah Three point two divides the sample space into subgroups: the decision output, accept or reject; process baseline characteristics such as site or equipment; the material or product; and the task, such as types and severity of defects. Everything downstream is done per subgroup: the test set must include them, the acceptance criteria may differ for them, and the size requirement applies to each.

09:23 Sam Acceptance criteria.

09:24 Sarah Second idea, section four: the pass mark is fixed before the exam. Test metrics are defined for the intended use; for a classifier the examples are a confusion matrix, sensitivity, specificity, accuracy, precision or F1 score. Acceptance criteria may differ by subgroup, a process expert is again responsible, and again they are approved before testing. Then four point three, titled No decrease. The acceptance criteria of a model should be, quoting, at least as high as the performance of the process it replaces. And the consequence, in the draft's own words: this implies that the performance should be known for the process which is to be replaced. It points to the revised Annex Eleven for that.

10:13 Sam Read the implication out loud, because it's the one from the cold open. If the acceptance criterion has to be at least the manual inspection's performance, you have to have measured the manual inspection. Most sites never have. They have a procedure, a training record, and an assumption. The draft turns the assumption into a number you must produce before the model is tested. That's a study, with its own test set, of people. Same logic as the parallel testing against the reference method in the Points to Consider and the orthogonal control in FDA's example: the manual process is the comparator.

Annex 22: test data and test execution

10:48 Sam Test data.

10:49 Sarah Section five, and the idea is that the test set is a designed object, not whatever was left over. Selection: representative of, and expanding, the full sample space; stratified; including all subgroups and the rare variations; with the criteria and rationale documented. Size: sufficient, for the whole set and for each subgroup, to calculate the metrics with adequate statistical confidence. Labelling: verified by a process that ensures, quoting, a very high degree of correctness; the examples are independent verification by multiple experts, validated equipment, or laboratory tests. Pre-processing pre-specified. Any exclusion documented and fully justified. And five point six: generating test data or labels, for example by generative AI, quoting, is not recommended and any use hereof should be fully justified.

11:47 Sam So no synthetic vials to pad out the reject class. And the test itself?

11:51 Sarah Section seven. It should show the model is, in the draft's phrase, generalising well, including detection of over- or underfitting. A test plan is approved before the test starts, containing the intended use, the pre-defined metrics and acceptance criteria, a reference to the test data, a test script, and how the metrics are calculated, with a process expert involved. Any deviation from the plan, any failure to meet the criteria, or any omission to use all of the test data is documented, investigated and fully justified. And everything is retained, including the actual test data, physical test objects, and the access-control and audit-trail records, like any other GMP documentation.

12:40 Sam Omission to use all the test data. That's a quiet clause with teeth. You can't drop the awkward subgroup on the day.

12:47 Sam Restate before independence, because that's the one that changes who does what. FDA wants a plan, then a report that matches it, sized to model risk, and nobody says how much is enough. Annex Twenty-two says start with an intended use owned by a process expert that characterises the whole input space. Fix criteria before testing, at least as high as the manual process, which you must therefore have measured. Design a stratified test set with verified labels and no synthetic data. Run it to an approved plan, investigate every deviation, keep everything. Now the part most data science teams have never been asked.

Independence is a people-and-access control

13:23 Sarah Section six, Test Data Independency. Six point one states the principle: technical or procedural controls to ensure that data used to test a model are not used during its development, training or validation. Two routes: capture the test data only after training is complete, or split them off from a complete pool before training starts. Six point two is the clause for the split route. It is, quoting, essential that employees involved in the development and training of the model have never had access to the test data. The test data are protected by access control and by an audit trail logging accesses and changes. And, quoting, there should be no copies of test data outside this repository.

14:13 Sam No copies. Not a spreadsheet on a laptop, not a notebook in a sandbox.

14:17 Sarah Six point three: record which data were used for testing, when, and how many times. Six point four, for physical objects: the objects used for the final test must not previously have been used to train or validate the model. Six point five, Staff independency: controls to prevent staff who have had access to the test data from training or validating the same model. Where an organisation cannot maintain that separation, such a person may work on training only together with a colleague who has not had that access, which the draft calls the four-eyes principle.

14:54 Sam Here's what that does. It converts data-science hygiene into an organisational control that quality can audit. Who had access is a question with an audit trail behind it. How many times the test set was used is a counter. Whether the vial in the final test ever passed under the camera during development is a chain-of-custody question. None of that is a modelling skill. It's the discipline the QC lab already has for reference standards, applied to a dataset.

15:21 Sarah FDA's step four asks for the same independence in a submission's voice. Test data, quoting, should be independent of the development data and should not be shown to the algorithm during training. The sponsor should specify how independence was achieved, and one of the examples is, quoting, data acquired using different batches or products. Any overlapping use of data must be explained and justified, and the reference method used to create the test data should be described, with a summary of its performance.

15:53 Sam Different batches or products. So for a process model, hold out whole batches, not random rows from every batch. A random split across a time series leaks: neighbouring points in the same batch are nearly the same point. The draft gives batches as the unit; the rest is my reading.

Explainability and confidence become evidence

16:12 Sam Explainability. This is the one that will surprise the vision vendors.

16:17 Sarah Section eight, two clauses. Eight point one, Feature attribution: during testing of models used in critical GMP applications, systems should capture and record, quoting, the features in the test data that have contributed to a particular classification or decision, and it gives rejection as the example. Where applicable, techniques such as SHAP values, Shapley Additive Explanations, or LIME, Local Interpretable Model-Agnostic Explanations, or visual tools like heat maps, should highlight the key factors. Eight point two, Feature justification: a review of those features should be part of the process for approving the test results.

17:00 Sam So the heat map isn't a demo. It's a record, produced during testing, reviewed as a condition of approving the test. And the reviewer has to be able to say the model rejected the vial for the meniscus and not for a scratch on the fixture.

17:14 Sarah Section nine, Confidence. Nine point one: when testing a model that predicts or classifies, the system should, where applicable, log the confidence score for each prediction. Nine point two: models should have an appropriate threshold, and if the confidence score is very low, it should be considered whether the model should, quoting, flag the outcome as undecided, rather than making potentially unreliable predictions or classifications.

17:41 Sam An undecided bin. Which means a downstream process for undecided vials, a person, a rate you monitor, and a number in the intended use for what you expect that rate to be. Most soft-sensor and vision pipelines I can picture do neither of these things. They emit a number. They don't say how sure they are, and nobody keeps the attribution.

Five FDA asks MSAT would not expect

18:02 Sam Now the FDA list, but grouped, because I don't want a checklist. Five things in step four a manufacturing science team would not expect.

18:11 Sarah First, statistics with error bars. The draft lists the metrics, area under the ROC curve, sensitivity, specificity, predictive values, precision, F1, and then the sentence that matters, quoting: all performance estimates should be provided with confidence intervals. Not a point accuracy; an interval, which needs a test set big enough to make it narrow. Second, independence explained, which we have covered. Third, provenance of anything you did not build: specify whether a pre-trained model was used, and if so, quoting, specify the dataset that was used for pre-training and how the pre-trained model was developed and/or obtained. The IMDRF principles say why: generative systems may employ foundation models not under the manufacturer's provenance.

19:00 Sam Which for the agent is the whole model.

19:02 Sarah Fourth, the human. In four B: if the context of use involves a human in the loop, quoting, ensure that the evaluation methods consider the performance of the human-AI team, rather than just the performance of the model in isolation. IMDRF principle seven says the same for devices, and lists the human factors: user skills, expertise, understanding of the model's limitations, and the potential for over-reliance.

19:31 Sam Flag for episode seven, and I'll keep it short. Evaluate the team means pick a human. Which human? The operator who runs the procedure at three in the morning, or the scientist who understands the model? Different tests, different results, and the draft doesn't say. If you evaluate with the scientist and deploy with the operator, you validated a team you don't run. Parked.

19:54 Sarah Parked. Fifth, software. Four A asks for, quoting, the quality assurance and control procedures of computer software (including its toolboxes and packages) and how version changes were tracked, and four B asks for the procedures for code verification. The model is software, and the libraries under it are software, with versions.

20:16 Sam Which means the Python environment is a configuration item. Pin it, record it, and the day a package updates is a change. Restate. FDA's structure is plan, execute, report, sized to risk. Annex Twenty-two fills the plan in: intended use owned by a process expert; criteria fixed in advance and no lower than the manual process; a designed test set; independence by access control and staff separation; attribution reviewed at approval; confidence logged with an undecided threshold. FDA adds confidence intervals, provenance, the human-AI team and software QA. Now the word. Why is all of this called assurance and not validation?

From validation to assurance: CSA and GMLP

21:04 Sarah Because FDA renamed it, on the device side, and the pharma side borrowed the name. The Computer Software Assurance guidance is from CDRH and CBER with CDER consulted: draft September twenty twenty-two, final the twenty-fourth of September twenty twenty-five, reissued the third of February twenty twenty-six. It is written for software in device production or a quality management system, so in pharma it applies by analogy, and the landscape records it as the document most cited as the model for AI validation in manufacturing. It defines assurance as, quoting, a risk-based approach for establishing and maintaining confidence that software is fit for its intended use, and it is least-burdensome: quoting, the burden of validation is no more than necessary to address the risk.

21:55 Sam What does it actually change?

21:57 Sarah Four moves. Identify the intended use. Decide whether a failure poses high process risk, meaning a quality problem that foreseeably compromises safety. Choose assurance activities commensurate with that: unscripted testing, scenario, error-guessing and exploratory, alongside scripted testing, and it says unscripted may be better suited even for high-risk features; plus leverage of the vendor's own validation work, of other controls in the process, and of the monitoring data the software generates in use. Then make the record: intended use, risk analysis, what was tested, issues found, a conclusion declaring acceptability, who and when, and approval. It names AI and machine learning tools, bots and cloud in its own scope.

22:52 Sam And the record's philosophy.

22:53 Sarah Two sentences. Documentation, quoting, need not include more evidence than necessary to show the software performs as intended for the risk identified. And FDA recommends, quoting, system logs, audit trails, and other data generated and maintained by the software, rather than paper or screenshots.

23:13 Sam Here's what I take from it. Validation, the old way, was proof that you had done the documents. Assurance is proof that you thought: here is the risk, here is what I did about it, here is what I found, here is why that's enough. And you're allowed to point at the log instead of printing it. Annex Twenty-two is the same posture with the content filled in for a model, and its own clause two point three sizes every activity to the risk to patient, product and data.

23:42 Sarah Two more sources supply the vocabulary and the checklist. The vocabulary is FDA's November twenty twenty-three device guidance on the credibility of computational modelling. It says of itself that it covers first-principles models and not machine learning; but FDA's twenty twenty-five draft says in its footnotes that its question of interest, context of use and model risk were informed by ASME V and V forty, and points to the twenty twenty-three guidance for decision consequence. Its definitions are the ones to hold. Credibility is trust, established through the collection of evidence, in a model's predictive capability for a context of use. Verification is about the code and the calculation.

24:28 Sam And validation, in that vocabulary.

24:30 Sarah Quoting: the process of determining the degree to which a model or a simulation is an accurate representation of the real world. It adds that the comparison must be against data independent of the data used to create the model, so calibration is not validation. Applicability is, quoting, the relevance of the validation activities to support the use of the computational model for a context of use. Credibility factors are the elements of the process, each with a gradation of rigour and a credibility goal chosen by model risk. Adequacy assessment then asks whether the evidence, taken together, is sufficient for the risk.

25:12 Sam So when someone says the model is validated, the credibility vocabulary asks three separate questions. Is the code right? Does it match reality on independent data? And does the test you ran resemble the use you have? The third is the one that gets skipped. It's applicability, and it's Annex Twenty-two's input sample space by another name.

25:32 Sarah The checklist is the ten principles of Good Machine Learning Practice, now IMDRF N eighty-eight, final, dated the twenty-seventh of January twenty twenty-five. As headings: intended use understood and multi-disciplinary expertise; good software engineering and security practices; representative datasets; training sets independent of test sets, with, quoting, the extent of external validation is proportionate to risk; reference standards fit for purpose; model choice tailored to the data and the intended use; the human-AI team assessed, not the device alone; testing under real conditions; clear information to users; and deployed models monitored, with retraining risks managed. They say patient where a plant would say process, and every one maps onto a clause we have read tonight.

The fill-volume camera's evidence package

26:24 Sam Now build one. The fill-volume camera from FDA's example, under Annex Twenty-two, clause by clause, and I'll play the site.

26:30 Sarah The system as FDA wrote it: Drug B, a parenteral in a multidose vial, fill volume a critical quality attribute, an AI visual system assessing fill level on every vial, with release testing still measuring fill volume on a sample per batch. High consequence, low influence, medium model risk, from episode two. That sets the dial. Section three, the intended use, owned by a process expert: classify each vial's fill level as within or outside specification from an image. The input sample space is every container format on the line, the full range of fill levels including the extremes, lighting across shifts and after lamp changes, glass variation between suppliers, label and cap positions, condensation and foam. Limitations: a partly occluded meniscus, a vial not seated in the fixture.

27:28 Sam Subgroups and criteria.

27:30 Sarah Under three point two: accept and reject; the line and the site; the container and the product; and the defect types, underfill, overfill, no fill, and their severity. Section four: sensitivity to a true out-of-specification vial, specificity, and the confusion matrix, per subgroup, with criteria stricter where the consequence is, because an underfill missed is a patient dose short. And four point three: those criteria must be at least the detection performance of the manual inspection the camera replaces. So the site runs a study of its inspectors on a known set of vials and writes that number down first.

28:12 Sam Which forces a conversation nobody had: what does the process actually present to the camera in a year? MSAT writes that, not the vendor.

28:20 Sarah Section five. The test set is designed: stratified across every subgroup, sized so each subgroup's sensitivity has a usable confidence interval, which for rare defects means deliberately making reject vials, physically, not synthetically, because five point six does not recommend generated data. Labels verified to a very high degree of correctness: each test vial's true fill measured gravimetrically or by validated equipment, not by another inspector's eye. Every vial excluded from the test set, and why, recorded. Section six: the test vials are physical objects, so six point four applies directly; they have never been imaged for training or tuning, they sit under access control with an audit trail, no copies of the images elsewhere, and the people who trained the model have never seen them.

29:15 Sam Let me push on one thing. The vendor trained the model. The vendor has the images. Does the site's independence clause reach the vendor?

29:23 Sarah Two point two says documentation for these activities should be available to, and reviewed by, the regulated user irrespective of whether the model was trained, validated and tested in-house or by a supplier. So the site owes the evidence that the vendor's training set and the site's test set are independent, and if it cannot get it, the route in six point one is to capture the test set after training, from vials the vendor has never had.

29:53 Sam Then the test, the heat maps and the threshold.

29:55 Sarah Section seven: an approved test plan, deviations investigated, everything retained. Section eight: the system records the image regions that drove each rejection, and a reviewer confirms at approval that the model is looking at the meniscus and not at the fixture or a reflection. Section nine: a confidence score per vial, a threshold below which the vial goes to undecided, and a procedure for undecided vials with the expected rate in the intended use.

30:26 Sam And then section ten from last time: change control, configuration control, performance and input-drift monitoring, so that when the lamps are changed the drift metric says so before the sensitivity does. That's the package. Notice how much of it is not about the neural network. It's about the vials, the inspectors, the access list and the reviewer.

Soft sensor, batch model and the agent

30:48 Sam Now the other three, and what changes. Soft sensor.

30:52 Sarah The soft sensor predicts a quality attribute from process signals, so its label is a laboratory measurement, and there is not one for every point in time. Five point three bites hardest: a very high degree of correctness means every test label is an assay result with its own method performance, which is the reference-method summary FDA asks for. Subgroups are campaigns, sites, equipment trains and product changeovers. Independence means holding out whole campaigns, which FDA's batches-or-products example supports. The no-decrease comparator is the sampling-and-assay process it stands in for. And feature attribution for a regression is which signals drove the prediction, with a process expert judging whether they are physically plausible.

31:39 Sam Batch monitoring model.

31:41 Sarah The multivariate model already lives in this world. Its explainability is the contribution plot the multivariate community has used for decades, which is feature attribution under another name, so eight point one is satisfied by recording and reviewing what already exists. Its test set is held-out batches, including known bad ones. Its confidence score is the distance statistic it already computes, and the undecided threshold is a control limit. The new asks are organisational: who chose the reference batches, whether the excluded batches are documented under five point five, and who saw the held-out batches before the model was fitted.

32:22 Sam And the agent. Take it slowly.

32:24 Sarah Under Annex Twenty-two the honest first sentence is that in critical use the agent never reaches the evidence sections, because section one says it should not be there. In non-critical use, with a qualified human owning every output, the draft says the principles may be considered where applicable. So apply them. Intended use: drafting a deviation investigation from batch records; the input sample space is every kind of deviation the site sees and every record format it reads. Subgroups: deviation type, product, severity, source system. Metric: the draft's classifier metrics do not fit a paragraph. The nearest any document comes is BioPhorum's June twenty twenty-six worked example of a deviation module, which proposes back-testing against closed deviations and CAPAs, cause-evidence explainability, mandatory investigator justification, hallucination flags and citation requirements.

33:25 Sam So the metric is agreement with an adjudicated set of historical investigations. And independence, for text?

33:32 Sarah Section six read literally: a held-out set of closed deviations, adjudicated by experts, in an access-controlled repository, no copies, and the people who tuned the prompts and the retrieval have never read them. Provenance, from FDA: which foundation model, pre-trained on what, obtained how, which for a commercial model the vendor may not fully answer. Explainability under section eight: a citation, which record, which line, for every factual claim in the draft, reviewed by the approver. Confidence under section nine: the model declines to draft where it cannot cite, an undecided for text. And the human-AI team under FDA: evaluate the engineer with the draft, not the draft alone.

34:24 Sam So the honest answer for the agent is an adjudicated set of historical deviations under lock and key, a citation-per-claim discipline, and a team evaluation. It's buildable. And in the EU it buys you a non-critical use with a person who owns every output, because the gate never opens further. That's not nothing. It's the first evidence package of its kind at most sites, and it will tell you more about your deviation records than about the model.

34:52 Sarah One divergence to note. BioPhorum's example grades deviation drafting as low consequence, operational efficiency, with a static locked model and human approval, which echoes FDA's operational-efficiency carve-out. Annex Twenty-two's question is not consequence but whether the application is critical, and a drafted investigation that shapes a batch disposition touches product quality and data integrity. That is BioPhorum's view set against a draft annex, and a site inspected in the EU reads the annex.

So what for biomanufacturing

35:25 Sam So what for biomanufacturing. Opinions labelled. First: the evidence discipline is knowable and mostly not new. Plan, criteria in advance, independent test, deviations investigated, records kept; the QC lab has done this for methods for decades. Three things are new, and they're organisational, not mathematical. Test-data independence is now an access list, an audit trail and a staffing rule. The no-decrease clause forces you to measure the manual process you're replacing, and most sites have never measured a human inspector or a sampling plan as a process with a performance. And explainability and confidence are records produced during testing and reviewed at approval, not a demo.

36:12 Sarah The documents place the first artefact the same way. Annex Twenty-two's section three intended use, owned by a process expert, is what an inspector reads first. FDA's step two context of use is what a reviewer reads first. They are the same document under two names.

36:29 Sam So start there. Second: write that document before you touch the model, and have MSAT write it, with QA and the data scientist in the room, because clause two point one of the annex puts them all there at algorithm selection. Third, for quality: treat the test set like a reference standard. A custodian, a location, an access log, a use counter, a rule that the people who train never see it, and a pair rule when you can't staff that. Here's my view, as the practitioner: if your data science team can't tell you who has seen the test set, you don't yet have evidence. You have results.

37:05 Sarah Where the teams feel it. MSAT and process science own the intended use, the subgroups, and the comparator study of the manual process. Data science owns the metrics, the confidence intervals, the attribution records and the software versioning. Quality owns test-data custody, the deviation investigations, and an approval that now includes the feature review. Regulatory CMC decides whether the report goes in the application or stays for inspection.

37:34 Sam And for the agent, one more, labelled. Build the adjudicated deviation set now, before anyone asks. It's the evidence for any language model you will ever run on those records, and building it will show you the state of your deviation history. It's the basement again. The agent is the one that tells you whether it's dry.

Recap and next time

37:52 Sam Three things. Sarah.

37:53 Sarah Evidence is scaled, never fixed. FDA's draft wants a credibility assessment plan, executed, then a report that matches it, all commensurate with model risk, filed or held for inspection. Nobody says how much is enough, and the plan-then-report discipline is itself the evidence.

38:13 Sam Second, mine: Annex Twenty-two is the recipe, and it's a draft an EU inspector already reads. Intended use owned by a process expert with the whole input space and its subgroups. Acceptance criteria fixed before testing and at least as good as the process you replace, which you must therefore measure. A designed test set with verified labels. Independence as an access-and-people control. Heat maps reviewed, and a confidence threshold with an undecided bin. The new work is organisational.

38:48 Sarah Third: validation became assurance because FDA's software assurance guidance sized the evidence to risk and let the record be the log; the credibility vocabulary, verification, validation, applicability and adequacy, came from the device modelling guidance; and the ten machine learning principles are the checklist that maps onto every clause. For the agent, none of it was written with text in mind, and the best available answer is an adjudicated set of historical deviations, held like a test set, with a citation for every claim.

39:23 Sam Next time: data. Are my data fit for use? Every framework tonight assumed the training set and the test set were trustworthy, relevant and reliable. Next episode reads what the data-integrity rulebook, the Annex Eleven revision and the AI texts actually require, and asks where that leaves a plant historian. Are we there yet? We know what to prove. Next we find out whether the data can prove it.