AI in Biopharma Manufacturing: Are We There Yet?
This is the script the AI voices read, so it matches the audio word for word; times are from the render. Researched, scripted and voiced by AI systems under Jack Prior's direction. Sam and Sarah are AI characters; nothing in the episode is Jack speaking, and none of it is a statement of his views or his employer's. Generative AI can be confidently wrong — check the sources. Corrections: jack@jackprior.ai.
00:00 Narrator This is Are We There Yet — a podcast on the evolution of AI in biopharma manufacturing, directed by Jack Prior. A word about how it's made. This episode was researched, scripted and voiced by AI systems. Jack sets the questions and the frames; the AI reads the documents and writes the conversation you're about to hear, between two AI characters: Sam, who plays a manufacturing-science practitioner, and Sarah, who has read the documents. Nothing you hear is Jack speaking, and none of it is a statement of his views or his employer's. Like any generative AI output, it can be wrong — confidently wrong, or missing a nuance — which are exactly the risks this industry is working to mitigate, and exactly what this season is about. Check the sources before you rely on anything. Corrections are welcome at jackprior dot A I. Now, the episode.
00:53 Sam A site sends a supplement to FDA. In it, a soft sensor on a production bioreactor, with a validation package two hundred pages thick: architecture, training data, accuracy, drift monitoring, the lot. The reviewer's first question back is one line. What decision does this model make, and what else do you rely on to make it? And the team realises that in two hundred pages, nobody wrote that sentence down. The process people know the answer. Quality knows a different answer. Regulatory wrote a third one in the cover letter.
01:22 Sarah And that one sentence is what FDA's framework asks for first, before it asks anything about the model. Sam, that's this episode.
01:30 Sam Context of use. Let's do it.
01:32 Sam Sarah, the question this time.
01:34 Sarah What is the one idea that appears in nearly every document in this rulebook, and where did it come from? The idea is context of use, and the answer is that it came from engineering standards for computer models of medical devices, and has now been adopted by FDA for AI in drugs and biologics and by ICH for model-informed drug development. It is the closest thing this field has to a shared grammar.
02:01 Sam Last time I said generative and agentic AI are why this season exists, and that they arrived faster than the rulebook. Context of use is where that shows first. For a soft sensor, the decision it informs is obvious. For an agent that reads a batch record and drafts an investigation, the decision is spread across a dozen small judgments and a human at the end. The framework we're reading tonight was built for the first kind of system. We're going to see how far it stretches toward the second.
02:31 Sam Which documents, and from where?
02:33 Sarah The anchor is FDA's draft guidance from January twenty twenty-five, Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products. Regulator guidance, layer two. It's issued by the drug centre together with the biologics, device and veterinary centres, so it speaks for the agency, not one office. Its maturity is draft: the comment docket closed in April twenty twenty-five with about a hundred and twenty comments, and it is still a draft as of this summer. Per last episode's house rule, that makes it the current expectation, not an option.
03:13 Sam And the voices.
03:14 Sarah Three. FDA's November twenty twenty-three guidance on assessing the credibility of computational modelling in medical device submissions, final, where the vocabulary was first written down for regulators. ICH M15 on model-informed drug development, adopted at Step Four in January twenty twenty-six, which arrives at the same terms from the clinical pharmacology side. And EMA's reflection paper on AI across the medicinal product lifecycle, final since September twenty twenty-four, which uses a different pair of axes. Last episode this was trend one, from models to AI with the same logic underneath. Tonight we read the logic itself.
03:59 Sam Start at the top. The team in my cold open had two hundred pages about the model. What does FDA want first?
04:05 Sarah Not the model. The decision. The draft lays out a seven-step framework, and the first three steps never mention architecture or accuracy. Step one is the question of interest: in the draft's words, the specific question, decision, or concern being addressed by the AI model. Step two is the context of use, which the draft defines as the specific role and scope of the AI model used to address a question of interest. Role is what the model does. Scope is whether it is the sole determinant or one input among others. The draft says the context of use should state whether other information will be used alongside the model output to answer the question.
04:51 Sam So scope is the sentence my team never wrote. What else do you rely on.
04:55 Sarah Exactly that. And it feeds step three, model risk, which the draft defines as a combination of two factors. Model influence: the contribution of the evidence derived from the AI model relative to other contributing evidence. And decision consequence: the significance of an adverse outcome resulting from an incorrect decision. The two are rated independently and combined in a matrix, the draft's Figure One, and model risk rises as either one rises.
05:24 Sam Independently. Say more about that, because I think it's where people go wrong.
05:29 Sarah A footnote in the draft is explicit. Decision consequence is the potential outcome of the overall decision, outside the scope of the AI model and irrespective of how modelling is used. So you rate consequence by asking how bad a wrong answer to the question would be, as if there were no model at all. Then you rate influence by asking how much the answer leans on the model. If you let the model's presence soften the consequence rating, you've counted the same mitigation twice.
05:59 Sam And the other trap: people hear model risk and think, is my model risky? Bad data, drift, black box.
06:06 Sarah The draft closes that door in one sentence. Model risk is, quoting, the possibility that the AI model output may lead to an incorrect decision that could result in an adverse outcome, and not risk intrinsic to the model. Data quality, drift and the model's limitations are real concerns, but they belong in step four, the credibility plan, as things you have to show. They are not inputs to the grade in step three. The grade is about the decision.
06:35 Sam That's a clean separation and I like it. My view: it's also why the framework is usable on a plant floor. You can grade a decision with the process owner and the quality lead in one meeting. You cannot grade a neural network in one meeting.
06:49 Sarah The draft's clinical example shows the grade at the other corner of the matrix. Drug A has a life-threatening adverse reaction, and a sponsor proposes an AI model to decide which trial participants can skip twenty-four-hour inpatient monitoring. The draft's scope sentence says only the AI model will be used to determine who is low risk. Sole determinant, so influence high. A wrongly classified participant could have the reaction at home, so consequence high. Model risk high. Same three steps as the factory example, opposite answer, and the difference is entirely in the scope sentence.
07:31 Sam Where did this vocabulary come from? It doesn't read like drug regulation.
07:35 Sarah It isn't, originally. In twenty eighteen the American Society of Mechanical Engineers published its standard V and V forty on the credibility of computational models for medical devices. Its sections two, three and four define question of interest, context of use, and model risk as influence times consequence. In November twenty twenty-three, FDA's device centre published a final guidance on assessing the credibility of computational modelling in device submissions. Its glossary defines context of use as a statement that defines the specific role and scope of the computational model used to address the question of interest, and it says it uses key concepts of the ASME standard but provides a more general framework.
08:20 Sam And the twenty twenty-five draft lifts it.
08:22 Sarah Openly. A footnote in the draft says the high-level concepts of steps one through three were informed by the ASME standard, sections two, three and four, and applies them to AI. Compare the definitions in the two documents and context of use, model influence, decision consequence and model risk are nearly word for word. The change is the noun: computational model becomes AI model.
08:48 Sam And ICH got there separately.
08:50 Sarah ICH M15, adopted in January twenty twenty-six, from model-informed drug development, where the models are mostly pharmacokinetic and exposure-response, though it names machine learning in its scope too. Its section two defines model influence as the intended weight of the model outcomes in decision-making considering the contribution of additional data or evidence. Consequence of wrong decision, rated on severity and likelihood together. Model risk, combining the two, and M15 adds that when the two ratings differ the risk may be driven by the more influential one.
09:27 Sam And the intrinsic-risk point?
09:29 Sarah M15 makes it too: model risk is not to be perceived as a risk intrinsic to modelling and simulation. Same sentence as FDA, different pen. And M15 gives the grade the same job: it says model risk is key for determining the requirements for model evaluation, and that evaluation should be commensurate with the model risk. Rate the decision, then scale the evidence.
09:53 Sam So a manufacturing scientist and a clinical pharmacologist now use the same four words for the same things.
10:00 Sarah With one caution. M15 has a fourth term, model impact, which means how far the modelling strategy departs from regulatory standards. It is not the same as model risk, and it is not the same as the older ICH idea of a high-impact model in manufacturing. That collision is next episode's subject, so tonight just note that impact in M15 is a different axis.
10:24 Sam Where we are. The framework starts with the decision: question of interest, then role and scope, then model risk as influence times consequence, with consequence rated as if the model didn't exist and model risk explicitly not the model's own risk. And the words come from device engineering via FDA's device centre, now shared with ICH. Now the example FDA chose.
10:50 Sarah FDA gives two worked examples and one is manufacturing. Drug B is a parenteral in a multidose vial. Fill volume is a critical quality attribute for release. The manufacturer proposes an AI-based visual analysis system doing one hundred percent automated assessment of fill level to identify deviations. Step one, the question of interest, in the draft's words: do vials of Drug B meet established fill volume specifications?
11:19 Sam Step two.
11:21 Sarah Role: the model analyses images of vials to determine whether a volume deviation has occurred. Scope: as part of release testing, independent verification of fill volume is performed on a representative sample of each batch, therefore the model is not the sole determinant for release. That sentence about the sample is the context of use doing its job.
11:43 Sam Step three.
11:44 Sarah Consequence first, as if there were no model. Volume is a critical quality attribute; releasing vials that don't meet it could lead to medication errors, the draft mentions an inability to withdraw the labelled content or pooling vials to make a dose. Decision consequence: high. Then influence. Because release testing measures fill volume on a sample, the draft says that reduces the model influence, rated low. High consequence, low influence, model risk medium.
12:14 Sam Now take the sample test away. Same camera, same accuracy, same everything.
12:18 Sarah Then the model is the sole determinant of release for fill volume. Influence becomes high. Consequence was never about the model, so it stays high. High and high is high model risk, and step four's credibility activities scale up accordingly. Nothing about the model changed. The decision architecture changed.
12:39 Sam This is the sentence I'd put on the wall. Keeping an orthogonal control in place is how you buy down model risk. It's the same move as parallel testing against the reference method in the old ICH implementation guide, and it's the same instinct behind the human-in-the-loop language everywhere else. And it means the cheapest validation strategy is often not a better model. It's a sample test you were going to run anyway.
13:02 Sarah The draft agrees in step seven. When credibility isn't sufficiently established for the model risk, the first of its five listed outcomes is that the sponsor may downgrade the model influence by incorporating additional types of evidence. Adding an orthogonal control is a design lever the framework hands you.
13:22 Sam Give me the rest of the seven quickly, and tell me why the examples stop.
13:26 Sarah Step four: develop a credibility assessment plan, describing the model, the data and the training, and how the model will be evaluated. Step five: execute it, ideally after discussing it with FDA. Step six: document the results in a credibility assessment report, including deviations from the plan. The report may go into a submission or be held and made available on request, for instance during an inspection. Step seven: determine whether the model is adequate for its context of use, with the five outcomes if it isn't.
14:03 Sam And why do the examples stop at three?
14:05 Sarah The draft says so directly: the two examples do not extend beyond step three because the credibility activities in step four are a general list, and the appropriate activities depend on the specifics of a programme that a hypothetical can't capture. So it will not tell you what enough evidence looks like for a medium-risk camera or a high-risk soft sensor. It tells you the grade, and the grade tells you how hard to work.
14:31 Sam We said that last time. It's deliberate, and it moves the burden to the site. Which is why the draft keeps saying talk to us.
14:38 Sarah Early engagement is its own section. For manufacturing the routes are the drug centre's Emerging Technology Program and the biologics centre's Advanced Technologies Team. And the wording matters. The draft says early engagement is highly encouraged before submitting a regulatory application or implementing an AI technology for drug or biological product manufacturing. Or implementing. That reaches an AI system a site deploys under GMP before any submission exists.
15:10 Sam That's a bigger reach than most people assume. A soft sensor that never appears in a filing is still, in FDA's mind, something to talk about before you switch it on. And a footnote reminds you the quality unit stays responsible regardless.
15:25 Sarah Yes. The footnote on the manufacturing example says the use of AI in production and process controls must be implemented in accordance with current good manufacturing practice, and that the quality control unit's responsibilities under parts two eleven point twenty-two and two eleven point sixty-eight apply. The framework sits inside GMP, not beside it.
15:47 Sam So far: decision first, then role and scope, then influence times consequence. The fill-volume camera is medium risk only because a sample test stays in place. Seven steps ending in a plan and a report, examples stopping at three by design, and FDA wanting a conversation before you implement, not just before you file. Now Europe, which counts differently.
16:11 Sarah EMA's reflection paper, final in September twenty twenty-four, doesn't fold everything into one model risk score. It uses two terms. High patient risk, for systems affecting patient safety. High regulatory impact, for cases where the impact on regulatory decision-making is substantial. They are rated separately, and the level of scrutiny, and whether EMA expects you to come and talk, depends on both. The paper also says the degree of risk depends not only on the AI technology and data quality, but also on the context of use and the degree of influence the AI technology exerts. So influence is in there, but as a factor rather than an axis.
16:55 Sam Does EMA want the conversation early, the way FDA does?
16:58 Sarah Yes, and it puts the responsibility squarely on the company. The paper says that if an AI system is expected to impact, even potentially, the benefit-risk balance of a medicinal product, early regulatory interaction is advised, and the level of scrutiny depends on the level of risk and regulatory impact. It also says it is the responsibility of the applicant or manufacturer to ensure that all algorithms, models, datasets and data processing pipelines are fit for purpose, and adds, quoting, these requirements may in some respects be stricter than what is considered standard practice in the field of data science.
17:40 Sam I'd underline that last line for every data scientist joining a plant. Give me the manufacturing case where the two axes come apart.
17:47 Sarah A soft sensor used for real-time process control can be high patient risk, because a wrong output can affect product quality, and low regulatory impact, because it never generates evidence for a submission. Turn it around: a model that generates the stability data behind a shelf-life claim may be low patient risk in operation and high regulatory impact, because a reviewer's decision rests on it. FDA's matrix would give both a single grade. EMA keeps them apart.
18:19 Sam And EMA on manufacturing specifically?
18:24 Sarah One paragraph. Section two point three point six says AI in manufacturing, including process design, scale-up, optimisation, in-process control and batch release, is expected to increase. Model development, performance assessment and lifecycle management should follow quality risk management principles, considering patient safety, data integrity and product quality, and ICH Q8, Q9 and Q10 should be considered, quoting, awaiting revision of current regulatory requirements and GMP standards. That last clause is EMA pointing at what became Annex Twenty-two.
19:03 Sam My read: for a global site the two systems are compatible. EMA's patient-risk axis is FDA's decision consequence wearing a different hat, and regulatory impact is the question FDA answers by asking whether the model is in scope at all. Write your context of use once and you can answer both.
19:22 Sam Now our four systems, through steps one to three. Soft sensor first, as feed-forward control. Sarah, run it.
19:30 Sarah Question of interest: should the feed rate on this bioreactor change now, and by how much? Role: the model predicts a cell density or titre from online signals and proposes the adjustment. Scope is where it splits. If the prediction goes straight to the controller, the model is the sole determinant of the feed decision, influence high. If it proposes and an operator confirms against limits, or the daily offline sample corrects it, influence drops to medium. Consequence: rated without the model. A wrong feed adjustment in a fed-batch culture affects product quality and possibly a batch, so medium to high depending on the process. Automated: high risk. Advisory with an orthogonal check: medium.
20:21 Sam So the question for the soft sensor is the fill-volume question. Where is the orthogonal control? The offline sample you already take, the operator limits, the downstream release testing. Name them in the scope sentence and the grade moves.
20:35 Sarah The draft asks for exactly that. In step one it says a variety of evidentiary sources may be used to answer the question of interest, including manufacturing process validation studies, and that these sources should be stated when describing the context of use in step two and are relevant when determining model influence in step three. An orthogonal control that isn't written into the scope doesn't count toward the grade.
21:02 Sarah Batch monitoring model next. Two contexts of use for the same model. As an operator alert, the question is: is this batch drifting from its expected trajectory? Role: flag a deviation for investigation. Scope: a human decides what to do; the batch is still released on its full release testing. Influence low to medium, consequence medium, because a missed alert delays a response but doesn't release product. Model risk low to medium. As release support, the question becomes: is this batch acceptable? If the trajectory fit substitutes for a release measurement, influence is high, consequence high, model risk high, and the credibility plan has to look like an analytical method validation.
21:49 Sam Same principal components, two different documents. That's the point of the frame: the model didn't change, its job did.
21:55 Sam Quick restate for the walkers. Three of our four are graded: soft sensor high or medium depending on whether the operator and the offline sample stay in the loop, batch model low as an alert and high as release support, camera medium with the sample test and high without. The agent is last, and it's the hard one.
22:14 Sarah Question of interest: what caused this deviation and what should the site do about it? Role: the agent reads the batch record and history and drafts an investigation with proposed actions. Scope: an engineer reviews, edits and owns the investigation; quality approves it. On paper that makes the engineer the determinant and the agent's influence low.
22:37 Sam On paper.
22:38 Sarah On paper. Two complications the documents don't resolve. First, scope. The draft's carve-out excludes AI used for operational efficiencies that don't impact drug quality. A deviation investigation decides root cause and corrective action on a quality event, so the agent is inside the framework the moment its draft shapes that outcome, which is its whole purpose.
23:04 Sam And the second?
23:05 Sarah Influence. The rating assumes the other evidence is genuinely independent. If the engineer reads the agent's draft and signs it, the draft is the evidence, and influence is high whatever the org chart says. The one place the draft speaks to this is in step four, where it says that if a human is in the loop, the evaluation should address the human-AI team and not the model alone. That's the evidence episode, but the frame comes from here.
23:34 Sam And consequence for the agent?
23:36 Sarah Rated without the agent: how bad is a wrong root cause or a wrong corrective action on a deviation? For a deviation touching a critical attribute, high. So the agent's model risk is medium if the human review is real and demonstrable, and high if it isn't. On EMA's axes it's high patient risk and low regulatory impact, unless the investigation ends up in a submission.
24:03 Sam Here's how I'd read that for a site. The agent's context of use has to be written so that the human is the determinant, and the site has to be able to show that's true, not assert it. That means the engineer re-derives the root cause from the record, not from the draft, at least on a sample, and the record shows it. If you can't design that review, the agent's influence is high and you should grade it that way and plan the evidence to match. My view, and the documents don't contradict it: the agent is the system where the scope sentence is hardest to write honestly, and the one where writing it honestly matters most.
24:42 Sam So what for biomanufacturing. Opinion section. First: define the context of use before anyone touches the model. Write the question of interest as one sentence, and the scope as one more: what the model does, and what else the decision relies on. Do that and half the validation argument writes itself, because the grade tells you how much evidence to gather and the scope tells you which orthogonal controls you're leaning on.
25:06 Sarah That's the draft's own ordering. Steps one to three before step four, and the draft says credibility activities should be commensurate with model risk and tailored to the specific context of use. The plan can't be written until the grade exists.
25:22 Sam Second: make quality, manufacturing science and regulatory sign the same sentence. In my cold open they each had their own. The reviewer will find the gap between them faster than you will. Third: treat orthogonal controls as a design decision, not a legacy. The sample test you were about to retire because the camera is so good is the thing keeping the camera at medium risk. Decide deliberately.
25:47 Sam Fourth, for the agent: write the scope sentence so a sceptical inspector would agree the human is the determinant, and build the review so it stays true under production pressure. If you can't, grade it high and say so. Fifth: talk to the regulator before you implement, not before you file. The draft says so for manufacturing, and the emerging technology programs exist for it. Sixth: data. The context of use tells you which data have to be fit for use. That's the basement again: you don't have to fix all of it, you have to fix the part under the decision.
26:24 Sarah And the draft defines fit for use. Relevant, which it glosses as including key data elements and sufficient data representative of the manufacturing process or operation. And reliable: accurate, complete, and traceable. For the camera that's the images and the sample results. For the agent it's every batch record it reads, which is why the agent asks more of the basement than the other three.
26:50 Sarah Who the documents put in the room: the draft names the quality unit as ultimately responsible, points the plan at the site's quality system, and routes early engagement through the emerging technology programs. EMA adds the regulatory-impact question, which is regulatory CMC's to answer. So quality, manufacturing science, and regulatory CMC, at the moment the sentence is written.
27:16 Sam Three things. Sarah, first.
27:17 Sarah The framework is about the decision, not the model. Question of interest, then role and scope, then model risk as influence times consequence, with consequence rated as if there were no model and model risk explicitly not the model's intrinsic risk.
27:33 Sam Second, mine: the orthogonal control is the lever. FDA's own camera is medium risk because a sample test stays in place, and high the day it's removed. Name your other evidence in the scope sentence and you've chosen your grade.
27:46 Sarah Third: the vocabulary is shared. It came from ASME's device standard through FDA's twenty twenty-three device guidance into the twenty twenty-five draft, and ICH M15 arrived at the same terms. EMA counts on two axes instead of one, but the context of use answers both.
28:05 Sam Next time: grading. Seven documents grade models seven ways, and impact doesn't mean impact. We'll anchor on the ICH implementation guide from twenty eleven, put all four of our systems in a tier, and move them. Are we there yet? We know what decision we're making. That's a start.