AI in Biopharma Manufacturing: Are We There Yet?

Transcript — Five Trends: How the Expectations Are Evolving

This is the script the AI voices read, so it matches the audio word for word; times are from the render. Researched, scripted and voiced by AI systems under Jack Prior's direction. Sam and Sarah are AI characters; nothing in the episode is Jack speaking, and none of it is a statement of his views or his employer's. Generative AI can be confidently wrong — check the sources. Corrections: jack@jackprior.ai.

00:00 Narrator This is Are We There Yet — a podcast on the evolution of AI in biopharma manufacturing, directed by Jack Prior. A word about how it's made. This episode was researched, scripted and voiced by AI systems. Jack sets the questions and the frames; the AI reads the documents and writes the conversation you're about to hear, between two AI characters: Sam, who plays a manufacturing-science practitioner, and Sarah, who has read the documents. Nothing you hear is Jack speaking, and none of it is a statement of his views or his employer's. Like any generative AI output, it can be wrong — confidently wrong, or missing a nuance — which are exactly the risks this industry is working to mitigate, and exactly what this season is about. Check the sources before you rely on anything. Corrections are welcome at jackprior dot A I. Now, the episode.

00:53 Sam Picture a Tuesday afternoon in a biologics plant. Two things on the slide. A soft sensor for a bioreactor, predicting from online signals an attribute the lab used to measure once a day. And below it, a pilot: an assistant built on a large language model that reads the batch record and drafts the deviation investigation for the engineer. The validation lead asks: which rule says we can do either of these? The room splits. The digital lead says there's no rule against it. Quality says the new European annex forbids the second one outright. Regulatory says FDA has a framework, but it's a draft and it never mentions language models. Everyone is right. Nobody has an answer.

01:39 Sarah And each of them is reading a different document, from a different layer of the rulebook, at a different stage of maturity. Sam, that's the whole problem this episode is about.

01:48 Sam Then let's start there.

Orientation: whose expectation, how settled

01:49 Sam Sarah, the question for this first episode, in one sentence.

01:53 Sarah Whose expectation is this, how settled is it, which way is it moving, and why do language models and agents make it urgent? Every later episode takes one idea, context of use, evidence, data, and reads it across the documents. This one is the map you need before any of that: who wrote each document, where it sits, and how far along it is.

02:18 Sam Let me take the urgency part first, because it's the reason this season exists. A soft sensor, a multivariate batch model, a camera on a filling line: those have been in plants for twenty years, under a rulebook written for them. What changed in the last two years is that generative and agentic AI walked into the plant faster than the rulebook could turn around. Agents that read batch records, draft investigations, propose actions to an engineer.

02:45 Sarah And you have a position on that, Sam.

02:47 Sam I do, and it's mine, offered for you to test. Agentic AI is the game changer. It's the thing that finally gives all that data work a use, and coding and data engineering as we've known them are heading into the rear-view mirror. And the rulebook is being rewritten in real time to catch up. This season reads that rewrite as it happens.

03:08 Sarah The rewrite is visible in the documents. Draft Annex Twenty-two's gate on generative models. FDA's silence on them. EMA's workshop in the summer of twenty twenty-six. NIST's generative AI profile. We'll place each of those tonight.

03:26 Sam Which documents, and from which layers?

03:28 Sarah Nearly all of them, briefly. The anchor this week isn't a regulatory text at all. It's the landscape map this season is built from: a working reference of about forty documents, with a lineage diagram and a timeline, last updated at the end of August twenty twenty-six. Two primary documents get read closely. The joint EMA and FDA Guiding Principles of Good AI Practice in Drug Development, published on the fourteenth of January twenty twenty-six: two pages, final, principles rather than guidance. And ICH's reflection paper on advanced pharmaceutical manufacturing, endorsed by the ICH Assembly on the eighth of October twenty twenty-five and published in March twenty twenty-six. That one proposes work; it isn't a guideline.

04:19 Sam And why five trends?

04:21 Sarah Because when you lay the documents out by date, they don't just accumulate; they shift in five directions. From models to AI. From model-agnostic to model-type-specific. From validation to assurance across a lifecycle. From regional to converging. And from silence on generative AI to a gap everyone can now see. Tonight each document gets named and placed, not summarised. The summaries are the rest of the season.

Four layers, four vocabularies

04:49 Sam Start with the map. Why do the three people in my Tuesday meeting talk straight past each other?

04:54 Sarah Because the rulebook has four layers, each with its own vocabulary; the landscape's first section lays them out. Layer one is law and regulation: in Europe the AI Act and EudraLex Volume Four, the GMP guide with its annexes; in the United States Title Twenty-one, parts two ten and two eleven, and Part Eleven on electronic records. Layer two is regulator guidance from FDA, EMA, the UK's MHRA, and PIC/S, the Pharmaceutical Inspection Co-operation Scheme. Layer three is harmonisation, meaning ICH, the International Council for Harmonisation, whose guidelines are adopted into regional guidance. Layer four is standards and industry practice: ISO, NIST, ASME, ISPE's GAMP guides, PDA, and consortia like BioPhorum.

05:52 Sam And the vocabulary problem?

05:54 Sarah Each layer names things differently. Critical GMP application is Annex Twenty-two language, from the law track. Context of use and model risk are credibility language, from FDA guidance and ICH. High-risk is a term of art in the AI Act. And impact means at least two different things inside ICH alone. So your quality lead is quoting a draft annex on the law track. Regulatory means an FDA guidance. And the digital lead is looking for a law and hasn't found one, because there isn't one yet.

06:27 Sam So they're all correct and all useless. My rule: the first job with any document in this space is to sort it by who wrote it and which layer it sits in, before you read a word of the content.

Non-binding is not a strategy

06:37 Sam Which brings up the phrase I want to kill this season. It's only guidance.

06:41 Sarah Here's what the documents say about themselves. FDA guidances describe the agency's current thinking, and a draft is current thinking on the day it's published. The European annexes are on the law track: once final, they are part of the GMP guide that inspectors inspect against. ICH guidelines are adopted regionally, which is how they acquire force. Standards are voluntary, but inspectors recognise them. The landscape's own layer table has a column headed binding, with a yes or a no in each row.

07:14 Sam And that column is the wrong headline. My opinion, clearly labelled: non-binding is a legal category. It is not a compliance strategy. If the reviewer has the January twenty twenty-five draft on their desk, they'll ask why your submission doesn't speak to it. If the inspector has read Annex Twenty-two in draft, they'll ask whether your application is critical. Nobody waits for the word final. So the axis that matters isn't binding or not. It's maturity, and who is going to ask.

07:42 Sarah The maturity axis runs in five steps, with an example of each in the landscape. Discussion paper: FDA's March twenty twenty-three paper on AI in drug manufacturing. Draft: FDA's January twenty twenty-five AI guidance and Annex Twenty-two, both drafts today. Final: EMA's reflection paper on AI in the medicinal product lifecycle, September twenty twenty-four. Adopted regionally: ICH M15, Step Four in January twenty twenty-six. Effective: the current Annex Eleven on computerised systems, in force since twenty eleven.

08:19 Sam And the smart move with a draft is two things at once. Comply early, because it's the current expectation. And comment through the docket, because the consultation is the only moment you can change the text. Annex Twenty-two drew a lot of comments.

08:34 Sarah About thirteen hundred, per the landscape, during a consultation that ran from the seventh of July to the seventh of October twenty twenty-five. That's an unusually large response for a GMP annex.

Trend one: from models to AI

08:46 Sam All right. Five trends. You've read across the whole pile. What's the first direction of travel?

08:51 Sarah Trend one: from models to AI, with the same logic underneath. It starts in twenty eleven with the ICH Quality Implementation Working Group's Points to Consider for Q8, Q9 and Q10. That document graded models by their role in assuring product quality. Low impact: models that support development. Medium impact: useful in assuring quality but not the sole indicator. High impact: the model's prediction is a significant indicator of quality, for example a chemometric model used for product assay. Validation and documentation scale with the grade.

09:29 Sam That's fifteen years old. Where does it go next?

09:32 Sarah To medical devices. In twenty eighteen, ASME published its standard V and V forty on the credibility of computational models. Its idea: model risk is model influence, how much the decision leans on the model, combined with decision consequence, how bad a wrong decision is, assessed for the specific question you are asking the model. FDA's November twenty twenty-three device guidance on the credibility of computational modelling adopted that logic.

10:03 Sam And then it jumps from devices to drugs.

10:05 Sarah In FDA's January twenty twenty-five draft on AI to support regulatory decision-making for drugs and biologics, which carries the same logic to AI as a seven-step credibility framework: question of interest, context of use, model risk, then plan, execute and document the credibility evidence, and finally judge whether the model is adequate for its context of use. And ICH M15, adopted in January twenty twenty-six, arrived at the same place from model-informed drug development: question of interest, context of use, model influence and consequence of a wrong decision, which together give model risk.

10:50 Sam So the AI part is almost incidental. Decide what the decision is, define the context of use, grade the risk, scale the evidence. That was true for a partial least squares model in twenty eleven.

11:01 Sarah ICH's reflection paper says almost exactly that. It credits the Points to Consider with the principle that a model's impact guides the extent of regulatory oversight, then says the document did not explicitly foresee these new types of AI models, and may not be best suited for AI models used as part of a dynamic control strategy and in continuous manufacturing. The logic holds; the coverage doesn't.

11:26 Sam And the open question under trend one?

11:28 Sarah How much evidence is enough. FDA's draft says credibility activities should be commensurate with model risk, and its worked examples stop at the risk grading, at step three. It does not describe what a sufficient evidence package looks like for a given grade. By design: proportionate, case by case, with early engagement encouraged.

11:50 Sam Deliberate, and expensive. When the regulator won't say what enough looks like, the site's own quality unit becomes the regulator. That's the evidence episode, later in the season.

11:59 Sam Where we are, for anyone walking. Four layers; sort a document by layer before you read it. Maturity and who will ask, not binding or not. And trend one: the logic for AI is the logic for models, context of use and risk-scaled evidence, from ICH in twenty eleven through devices to FDA's draft and M15. Trend two.

Trend two: Annex 22 breaks ranks

12:20 Sarah From model-agnostic to model-type-specific. FDA and EMA both stayed agnostic. FDA's draft doesn't care what kind of model you use; it grades the decision. EMA's reflection paper covers every kind of machine-learning model, warns that generative language models produce plausible but erroneous output and wants close human supervision where they draft documents, and scales scrutiny by risk. Then in July twenty twenty-five the European Commission published draft Annex Twenty-two on artificial intelligence, a new GMP annex written specifically about AI. Per the landscape it was drafted by EMA's GMP and GDP Inspectors Working Group and co-published with PIC/S. Its scope is deliberately narrow: static, deterministic machine learning models in critical GMP applications.

13:21 Sam And what happens to everything outside that scope?

13:24 Sarah This is the part that broke ranks. Dynamic models, the ones that keep learning after deployment, and generative AI including large language models, are not simply left out. The draft says such models should not be used in critical GMP applications. Generative AI in non-critical applications is allowed with a human in the loop. Those thirteen hundred comments pushed back, and EMA held a workshop at the end of June twenty twenty-six on adaptive and generative models. As far as the public record shows, the draft text has not changed since. The landscape has the final text targeted to reach the Commission around the end of twenty twenty-six; the effective date isn't set, and as I understand it people are planning for twenty twenty-seven.

14:12 Sam And FDA is going the other way.

14:14 Sarah FDA's draft describes AI models as data-driven and, in its words, self-evolving, capable of autonomously adapting without any human intervention. It tells sponsors to anticipate model-directed changes and manage them under lifecycle maintenance. So one authority anticipates a learning model and asks you to govern it. The other says don't put one in a critical application. That's the live divergence.

14:39 Sam For a global site the consequence is simple. If a European inspector will walk your floor, you build static. Freeze the model, version it, retrain under change control, learning loop outside the GMP boundary. My view: for a soft sensor a site should do that anyway. A model that retrains itself nightly is a model nobody can explain a month later. Whether the prohibition is wise is episode four.

Trend three: validation to lifecycle

15:05 Sam Trend three.

15:07 Sarah From validation to assurance and lifecycle. Three documents carry it. First, FDA's Computer Software Assurance guidance for production and quality management system software, from the device and biologics centres. Drafted in twenty twenty-two, final on the twenty-fourth of September twenty twenty-five, revised February twenty twenty-six. It reframes computerised system validation as risk-based assurance: test where the risk is, lean on the vendor's work where you can. It's device guidance, but pharma uses it by analogy.

15:42 Sam Second?

15:43 Sarah The Annex Eleven revision, drafted in the same July twenty twenty-five consultation as Annex Twenty-two. The current Annex Eleven is short. The revision runs to seventeen chapters: risk-based validation, supplier oversight, audit trail, identity and access management, cybersecurity, periodic review. Annex Twenty-two hangs off it, and a revised Chapter Four on documentation sits under both. Read the three as one rewrite of computerised-systems GMP. Finals are expected late twenty twenty-six, with implementation, as I understand it, in early twenty twenty-seven.

16:23 Sam And the third is FDA again.

16:25 Sarah FDA's January twenty twenty-five draft, in its lifecycle section. It asks for a life cycle maintenance plan with performance metrics, a risk-based monitoring frequency, and triggers for retesting the model. The detailed plan lives in the manufacturing site's pharmaceutical quality system, and a summary goes into the marketing application for any product- or process-specific model. Changes to the model, and changes in manufacturing that could affect the model, go through the site's change management system. And then the draft says sponsors may also use the tools in ICH Q12, naming established conditions and comparability protocols.

17:04 Sam Established conditions — for anyone who hasn't lived through a Q12 filing, say what that means.

17:09 Sarah Q12's term for the parts of your approved dossier that are legally binding: the process parameters, specifications and controls the regulator signed off on, where a change is a regulatory submission with a reporting category — not a note in the site's change control. Everything else in the filing is supportive information. So if elements of a model become established conditions, retraining it is a regulatory change, not an IT ticket. The draft opens that door without walking through it. It doesn't say which elements of a model would be established conditions, or which reporting category a retrain would fall into. That's the change episode.

17:50 Sam The headline of trend three, in my words: you own the model's whole life. Not a validation report in a binder. A living thing with metrics, a monitoring schedule and retest triggers, inside the quality system. A model that's validated and then left alone decays, and nothing in these documents will save it.

18:09 Sam So far, three trends. Models to AI with the same logic. Model-agnostic to Annex Twenty-two's model-type line, with FDA on the other side. And validation to assurance, where you own the lifecycle. Trend four is the hopeful one.

Trend four: the joint principles

18:23 Sarah From regional to converging. On the fourteenth of January twenty twenty-six, EMA and FDA jointly published Guiding Principles of Good AI Practice in Drug Development. It calls itself initial collaborative work; per the landscape, it's the first time the two regulators have put their names to one set of AI principles for medicines. Two pages, ten principles. Its scope is AI used to generate or analyse evidence across the drug product life cycle, and it lists the phases: nonclinical, clinical, post-marketing, and manufacturing. Manufacturing is named explicitly.

19:05 Sam Give me the ten, quickly.

19:06 Sarah Human-centric by design. A risk-based approach, with validation and oversight proportionate to, in its words, the context of use and determined model risk. Adherence to standards, including GxP. A clear context of use, defined as role and scope. Multidisciplinary expertise across the lifecycle. Data governance and documentation, with provenance and processing steps traceable in line with GxP. Model design and development practices, using data that are fit for use. Risk-based performance assessment of, quoting again, the complete system including human-AI interactions. Life cycle management under a risk-based quality system, with scheduled monitoring and periodic re-evaluation, and it names data drift. And clear, essential information for users in plain language.

20:02 Sam Those are FDA's words. Context of use, model risk.

20:05 Sarah And EMA's: human-centric and data governance come from its reflection paper. It reads like the two frameworks shaking hands. On maturity, be precise: it's final, but it describes itself as laying a foundation for good practice and identifying areas where regulators and standards bodies could collaborate, which may inform regulatory policies and guidelines. A direction of travel, published jointly, not a guidance with a docket.

20:33 Sam And ICH is where this would become permanent.

ICH's reflection paper

20:36 Sarah That's the second document. ICH's reflection paper, endorsed by the Assembly on the eighth of October twenty twenty-five, proposes guideline work on three topics: process modelling, explicitly including AI-based models; continuous process verification, in Q eight's sense of validating a process by monitoring it rather than on a fixed number of batches, not to be confused with the continued process verification every commercial process already owes; and decentralised or distributed manufacturing. Its preferred first step is a new ICH guideline on process models, with a revision of the Points to Consider to link model risk to intended use and decision consequence. Its open questions include what release testing is expected when a model controls the process, and what lifecycle maintenance means for, in its words, frequently updating process models, for example AI models that learn and self-adjust.

21:40 Sam So the most likely permanent home for harmonised AI in CMC is a guideline that hasn't been formally proposed yet.

21:47 Sarah Correct. The paper is endorsed; no new topic has been adopted. Meanwhile industry is stitching the pieces together itself. BioPhorum published AI risk guidance for the pharmaceutical industry in June twenty twenty-six, subtitled harmonizing frameworks for practical implementation. It walks across the AI Act, EMA's paper, FDA's draft, Annex Twenty-two, NIST, ISO and ICH Q9, and proposes one harmonised GxP AI risk framework, which it says aligns with the joint principles. Layer four doing the integration the upper layers haven't finished.

22:27 Sam My read: the convergence is real in vocabulary and not yet in rules. Still good news for a site. Write your model justification in context of use and model risk and the document travels; Rockville or Amsterdam, they'll recognise it.

Trend five: silence on generative AI

22:43 Sam Trend five. Generative AI. This is the one the Tuesday meeting was really about.

22:47 Sarah From silence to a gap everyone can see. FDA's January twenty twenty-five draft never uses the words generative AI or large language model. Its only handle is a scope carve-out. It says the guidance does not address AI used for operational efficiencies, and gives examples: internal workflows, resource allocation, drafting or writing a regulatory submission, provided those uses do not impact patient safety, drug quality, or the reliability of study results. That carve-out holds right up to the moment a language model's output shapes a quality decision. Then you're back inside the framework with a model nobody has described.

23:29 Sam And Europe?

23:30 Sarah Annex Twenty-two says generative AI should not be used in critical GMP applications, and requires a human in the loop for non-critical ones. EMA's reflection paper warns that generative language models produce plausible but erroneous output. The risks themselves are named outside pharma, in NIST's Generative AI Profile from July twenty twenty-four: confabulation, information integrity, and content provenance as a control. And one thing is under-addressed on every layer, per the landscape's open-questions list: third-party dependency. When a vendor updates the foundation model under your validated application, outside your change control, no document says what that does to your validated state. The nearest hook is supplier oversight in the Annex Eleven revision.

The four recurring systems

24:23 Sam Hold that thought. Five trends, then. Models to AI. Agnostic to model-type. Validation to lifecycle. Regional to converging. Silence on generative AI to a visible gap. Now the four systems that will carry every episode this season.

24:39 Sam First, a bioreactor soft sensor: it takes the signals you already have online, pH, dissolved oxygen, capacitance, off-gas, and predicts something you can't measure continuously, a cell density or a titre, to inform a feed or harvest decision. Second, a multivariate batch monitoring model, MVDA or MSPC: principal components over a batch trajectory, flagging when a run drifts from the golden profile so an operator looks. Third, FDA's own example from the twenty twenty-five draft: a camera system that reads vial images for a fill-volume deviation. Fourth, the agentic assistant: an agent built on a large language model that reads the batch record, drafts the deviation investigation, and proposes actions to an engineer.

25:24 Sarah And in FDA's example, release testing still measures fill volume on a sample of every batch, so the model isn't the sole determinant of release. FDA grades the consequence high, the influence low, and the model risk medium.

25:38 Sam Which trend bites each of the four hardest?

25:41 Sarah The fill-volume system is trend one: the whole example exists to teach context of use and the value of an orthogonal control; remove the sample test and the same model re-grades. The batch monitoring model is trend three, lifecycle: the golden profile drifts with every process change, so the question is who owns the model next year. The soft sensor is trend two, model type: the temptation is to keep retraining it, and the moment it retrains itself in a critical application, Annex Twenty-two says stop. And the agent is trend five, silence: it's the one system on the list that no document describes.

The agent on the maturity axis

26:26 Sam So take the agent down the maturity axis. Which documents touch it, at which stage, and what does each actually say about it?

26:32 Sarah Start with what is effective now. Parts two ten and two eleven, and the EU GMP guide, make deviation investigation and batch record review the quality unit's responsibility. Part Eleven, the current Annex Eleven, and FDA's twenty eighteen data integrity questions and answers govern the records the agent reads and any record it writes. None of those texts contemplates a machine drafting the investigation. None forbids a draft, so long as a qualified person reviews and owns it. The one text that draws the line directly is industry: BioPhorum's June twenty twenty-six risk guidance says AI systems must not directly execute electronic signatures.

27:19 Sam So the effective layer says nothing about the agent and everything about the records around it.

27:24 Sarah Yes. Next, the final documents. EMA's reflection paper warns that generative language models produce plausible but erroneous output and wants close human supervision. The joint principles apply in full: a clear context of use, human-centric design, and principle eight, performance assessment of the complete system including human-AI interactions. NIST's generative AI profile names the failure modes that assessment would test for. And on the practice layer, ISPE's GAMP Guide on artificial intelligence, from July twenty twenty-five, is the industry reference the landscape lists for AI across a GxP lifecycle.

28:06 Sam Drafts?

28:07 Sarah Two, and they pull in different directions. Annex Twenty-two says generative AI should not be used in a critical GMP application, and requires a human in the loop for non-critical ones. So the entire question becomes whether drafting a deviation investigation for human review is a critical application. The draft grades the application as critical or not, and separately rules certain model types out of critical use; the landscape's crosswalk names this exact case as the open question.

28:39 Sam And FDA?

28:41 Sarah Silent. The agent lives in the operational-efficiency carve-out until its draft shapes a quality decision, and a deviation investigation is a quality decision. At that moment the agent is inside a seven-step framework that has no example for it. And the Annex Eleven revision adds supplier oversight, which matters because the foundation model under the agent is a supplier.

29:05 Sam And the bottom of the axis?

29:06 Sarah Discussion and plans. CDER's promised guidance on AI and machine learning quality considerations in pharmaceutical manufacturing, on the twenty twenty-six agenda and not yet published. A possible EMA follow-up after the workshop; nothing announced. And ICH's reflection paper, which is about process models. The agent is not a process model at all, so even the likeliest harmonisation home has no room for it yet.

29:33 Sam So for the agent the honest answer is: no document yet says. A site has to reason it out anyway, and this is how I'd do it. Three questions. First, criticality of the decision it touches. Is the agent's output a draft an engineer owns, or the basis for a disposition decision nobody re-derives? Design for the first and be able to prove it.

29:55 Sam Second, a human in the loop that is real, not a signature on a page the agent wrote. The joint principles say assess the complete system including the human-AI interaction; the test is whether you can show the engineer caught the agent's mistakes. Third, data integrity of what it reads. The agent is only as good as the batch record, and what it writes becomes a GMP record with an authorship question attached. Pin the model version, treat every vendor update as a change, keep the signature human. None of that is in a document. All of it will be asked about.

30:31 Sarah What the documents support, Sam: every one of those three questions maps onto existing text. Criticality is Annex Twenty-two's own axis. Human oversight is the joint principles and EMA's paper. Data integrity is the effective layer, in force today. The judgment about where the agent's use falls is the part no document makes for you.

30:53 Sam Which brings us to what's coming. Give me the watchlist as it stands, and hedge where you should.

30:58 Sarah Six items, from the landscape's watchlist. CDER's AI and machine learning quality considerations guidance, on the twenty twenty-six agenda; FDA's first CMC-specific AI text. Final Annex Twenty-two, Annex Eleven and Chapter Four, late twenty twenty-six, effective around twenty twenty-seven, dates not fixed. A possible EMA follow-up on adaptive and generative AI; nothing announced. FDA's twenty twenty-five draft going final; no date. An ICH new-topic proposal on process models; none yet. And the EU AI Act, in force since August twenty twenty-four, its high-risk obligations pushed back by an amending regulation in July twenty twenty-six to December twenty twenty-seven and August twenty twenty-eight, per the landscape. Most manufacturing AI isn't high-risk under the Act unless it's a safety component of a regulated product, a medical device or a machine, that needs third-party conformity assessment.

So what for biomanufacturing

32:02 Sam So what for biomanufacturing. This is the opinion section. First: stop sorting documents into must and optional. Sort them by how settled they are and by who is going to ask. A draft annex an inspector has read is more real on your floor than a final standard nobody inspects against. Second: make the watchlist a habit, not a project. Someone owns the docket calendar, and when the next draft opens, the site comments, because that's the only moment the text can move and because writing the comment forces you to work out what you think.

32:35 Sam Third, for the agent specifically: decide now which uses you will argue are non-critical, write the argument down, and be ready to defend it to an inspector who has read Annex Twenty-two. If you can't defend it, the use is critical, and under the current draft that means the agent doesn't go there. Fourth, the one I care most about: data readiness is the prerequisite for every one of the five trends. The credibility framework assumes fit-for-use data. Annex Eleven and Chapter Four are about data integrity before they're about anything else.

33:09 Sarah Principle six in the joint document: data source provenance, processing steps and analytical decisions documented in a detailed, traceable and verifiable manner. That is the ALCOA expectation FDA's data integrity guidance has carried since twenty eighteen, and the ALCOA-plus version in PIC/S's data integrity guide since twenty twenty-one.

33:33 Sam That's the basement under the house. Plants have spent years pouring the foundation, historians, data lakes, contextualisation, and most of it is still a very good basement with no house on top. And here's the twist. It won't be the soft sensor that tells you whether the basement is dry; it reads six clean signals. The agent reads everything: the batch record, the deviation history, the trends, the SOP. The first week an agent runs on a plant's records, it will find every gap and every undocumented decision in them. That's not a reason to hold back. It's the reason to start.

34:08 Sarah The joint principles back the last part of that: principle five asks for multidisciplinary expertise covering both the AI technology and its context of use, and FDA's draft puts the lifecycle plan inside the site's quality system, so quality owns the monitoring, not the team that built the model.

34:27 Sam Which says who has to be in the room. Manufacturing science, because they own the question of interest. Quality, because they own the lifecycle plan and change control. Regulatory CMC, because the summary goes in the application. IT and data, because provenance and access control are theirs. And the vendor, because trend five says their update is your problem.

Recap and next episode

34:49 Sam Three things to take away. Sarah, the first.

34:51 Sarah The rulebook is four layers with four vocabularies: law, regulator guidance, ICH harmonisation, and standards and practice. Place a document by who wrote it and where it sits before you read it.

35:06 Sam Second, mine: non-binding is a legal category, not a compliance strategy. Sort by maturity, discussion paper, draft, final, adopted, effective, and by who will ask. Build to the drafts and comment through the docket.

35:20 Sarah Third: the five trends. From models to AI on the same credibility logic. From model-agnostic to Annex Twenty-two's line on model type, with FDA on the other side. From validation to owning the lifecycle. From regional to converging, with the joint principles and ICH's reflection paper marking the path. And from silence on generative AI to a gap the documents now acknowledge, with the agent standing in it.

35:49 Sam Next time we take the one idea that shows up in nearly every one of these documents: context of use. What decision is the model making? We'll anchor on FDA's January twenty twenty-five draft and walk all four of our systems through its first three steps, including the fill-volume example with the sample test taken away, and the agent, for which the framework was never written. Are we there yet? Not yet. But we know where the map is.