AI in Biopharma Manufacturing: Are We There Yet?

Transcript — Grading: How Much Does This Model Matter?

This is the script the AI voices read, so it matches the audio word for word; times are from the render. Researched, scripted and voiced by AI systems under Jack Prior's direction. Sam and Sarah are AI characters; nothing in the episode is Jack speaking, and none of it is a statement of his views or his employer's. Generative AI can be confidently wrong — check the sources. Corrections: jack@jackprior.ai.

00:00 Narrator This is Are We There Yet — a podcast on the evolution of AI in biopharma manufacturing, directed by Jack Prior. A word about how it's made. This episode was researched, scripted and voiced by AI systems. Jack sets the questions and the frames; the AI reads the documents and writes the conversation you're about to hear, between two AI characters: Sam, who plays a manufacturing-science practitioner, and Sarah, who has read the documents. Nothing you hear is Jack speaking, and none of it is a statement of his views or his employer's. Like any generative AI output, it can be wrong — confidently wrong, or missing a nuance — which are exactly the risks this industry is working to mitigate, and exactly what this season is about. Check the sources before you rely on anything. Corrections are welcome at jackprior dot A I. Now, the episode.

00:53 Sam A global site is doing the risk assessment for one batch monitoring model. One model, one spreadsheet, one column headed risk. The validation lead, trained on ICH, writes medium impact. The European quality lead writes critical application. The regulatory team, who've read FDA's draft, write medium model risk. And the digital team, using a template built on the new industry guidance, write low model influence. Four labels, all defensible, and the site's governance board asks the only reasonable question: is this thing risky or not? Meanwhile legal has sent a memo saying the deviation-drafting agent is not high-risk under the AI Act, and half the room thinks that settles it.

01:36 Sarah Seven documents grade models, Sam, and they grade different things: the model, the decision, the application, the product, or the feature. None of the four labels is wrong. They just aren't answers to the same question.

01:50 Sam So tonight we build the crosswalk.

Orientation: seven gradings

01:52 Sam Sarah, the question.

01:53 Sarah Seven documents grade models seven ways. How do they map onto each other? Last episode gave us model risk as influence times consequence, from FDA's draft and ICH M15. Tonight we set that beside the older ICH grading, Europe's critical-or-not test, the AI Act's tiers, and the industry framework from BioPhorum, and end with four tiers that this season will use from here on.

02:20 Sam And why it matters now, briefly, because I said my piece in episode one. For the classical models, grading is housekeeping. For the agent it's the whole fight. The European draft has no middle tier: an application is critical or it isn't, and if it's critical, a language model shouldn't be in it, however small its influence. So where you place the agent on the grading decides whether it exists at all on an EU-inspected site. That's not housekeeping.

02:48 Sam Documents.

02:49 Sarah The anchor is the ICH Quality Implementation Working Group's Points to Consider for Q8, Q9 and Q10, from twenty eleven. Layer three, harmonisation. It is an ICH-endorsed implementation guide, final, and it has a section on models with a low, medium, high grading that predates AI in the plant by a decade. Be clear on one thing: it is about models generally, not AI. That is the point of reading it. Beside it we read the crosswalk file this season is built on, which lays the seven gradings side by side.

03:24 Sam And the voices.

03:25 Sarah Four. ICH Q9 revision one on quality risk management, final in January twenty twenty-three, for the principle that formality scales with risk. Draft Annex Twenty-two, from July twenty twenty-five, for the critical application test and the model-type gate. The EU AI Act, regulation twenty twenty-four slash sixteen eighty-nine, in force since August twenty twenty-four, for Article Six and Annex Three. And BioPhorum's AI risk guidance, dated May and released in June twenty twenty-six, for the industry framework that redefines one of FDA's terms.

ICH 2011: low, medium, high impact

04:01 Sam Start with the oldest. What did ICH say about models in twenty eleven?

04:05 Sarah Section five of the Points to Consider opens by defining a model as a simplified representation of a system using mathematical terms, and says models can be used at every stage of development and manufacturing. Then section five point one grades them. The key sentence: for regulatory submissions, an important factor is the model's contribution in assuring the quality of the product, and the level of oversight should be commensurate with the level of risk associated with the use of the specific model.

04:37 Sam And the grades.

04:38 Sarah Three. Low-impact models support development, formulation optimisation is the example. Medium-impact models, quoting, can be useful in assuring quality of the product but are not the sole indicators of product quality; most design space models and many in-process controls. High-impact models, where the prediction is a significant indicator of quality of the product; a chemometric model for product assay, a surrogate model for dissolution.

05:08 Sam Sole indicator. That's our scope sentence from last episode, fourteen years early.

05:12 Sarah And the document's own examples make the same move FDA made with the fill-volume camera. It says a multivariate statistical process control model used for continuous process verification along with a traditional method for release testing would likely be medium impact. Continuous in Q eight's sense, validating the process by monitoring it, not to be confused with the continued process verification every commercial process owes anyway. The same model used to support a surrogate for release testing in a real-time release approach would likely be high impact. Same model, different role, different grade. A calibration model behind a near-infrared method is high impact if the method is used for release. A feed-forward model adjusting compression parameters from incoming material attributes is given as medium.

05:59 Sam So the traditional release test sitting alongside the model is what keeps it at medium. That's the orthogonal control, in ICH's words, in twenty eleven.

06:09 Sarah In its words: along with a traditional method for release testing. And the grade drives the work. Section five point three lists validation and verification elements it calls appropriate for high-impact models: acceptance criteria tied to the model's purpose, comparing prediction accuracy against the reference method, validation on an external data set from batches not used to build the model, parallel testing with the reference method at implementation and repeated through the lifecycle, and verification at commercial scale. For medium and low-impact models it says the applicability of those elements can be considered case by case.

06:50 Sam And documentation?

06:52 Sarah Scales the same way. Section five point two lists nine development steps, and it's worth noticing the first: defining the purpose of the model, before choosing the modelling approach or the variables. That is the question of interest, in twenty eleven. Step eight is evaluating the effect of prediction uncertainty on product quality and reducing the residual risk through the control strategy, which it says applies to high and medium-impact models. And step nine is documenting the model and planning its verification and update through the lifecycle, with the level of documentation dependent on the impact of the model.

Q9: formality scales with risk

07:32 Sam Which is Q9's principle applied to models.

07:34 Sarah Q9 revision one states it twice. As a primary principle: the level of effort, formality and documentation of the quality risk management process should be commensurate with the level of risk. And in its section on formality: formality is not binary but a continuum, and how much to apply depends on uncertainty, importance and complexity. One line worth carrying: resource constraints should not be used to justify the use of lower levels of formality.

08:06 Sam My view: this is the spine. If a site wants one grading to hang a validation argument on, it should be this one. It's ICH-endorsed, it's been adopted for fourteen years, and no inspector will argue with it. Everything else we read tonight either descends from it or has to be reconciled with it.

Impact is not impact

08:25 Sam Now the trap. Impact.

08:27 Sarah Two ICH documents use the word for different things. In the Points to Consider, a high-impact model is one whose prediction is a significant indicator of product quality. In ICH M15, from January twenty twenty-six, model impact means the extent to which the proposed modelling strategy varies from regulatory standards, or from expectations where no standard exists. That is about novelty of approach, not weight in the decision. M15's equivalent of the twenty eleven idea is model risk, influence times consequence, which we read last episode. So a clinical pharmacologist saying high model impact means an unusual modelling strategy. A manufacturing scientist saying high-impact model means a release surrogate.

09:15 Sam Practical rule: anyone who bridges CMC and clinical pharmacology says which one they mean, every time. And in a manufacturing document, if you want the M15 concept, say model risk.

09:26 Sam Where we are. ICH graded models in twenty eleven by their role in assuring quality: low, medium, high, sole indicator or not, with the release test alongside the model as the thing that holds it at medium. Rigour scales with the grade, per Q9. And impact in M15 is a different word. Now Europe, which doesn't grade the model at all.

Annex 22 grades the application

09:49 Sarah Draft Annex Twenty-two grades the application. Its scope, section one: it applies to computerised systems in manufacturing where AI models are used in, quoting, critical applications with direct impact on patient safety, product quality or data integrity. Critical or not; that is the whole scale.

10:09 Sam And then the gate.

10:10 Sarah A model-type gate on top. Static models only, meaning models that don't adapt during use. Deterministic only, meaning identical inputs give identical outputs. Dynamic models and probabilistic models are not covered and, it says, should not be used in critical GMP applications. And then, quoting, the document does not apply to Generative AI and Large Language Models, and such models should not be used in critical GMP applications.

10:40 Sam So there's no influence axis. A model that gives an operator a hint in a critical step is graded the same as a model that releases the batch.

10:48 Sarah Correct, and that's the gap the crosswalk calls G one: no middle tier. Every other framework has a gradient. Annex Twenty-two is binary plus a prohibition. There is no tier for critical application, low model influence, strong intervening controls, which describes most real advisory deployments; the nearest the draft comes is its human-in-the-loop clause. A low-influence advisory language model in a critical step is still should not be used under the current draft.

11:20 Sam Does it say anything about the human at all?

11:22 Sarah One clause, section three point three, on the human in the loop: where a model gives input to a decision made by a human operator, and the effort to test the model has been diminished because of that, the intended use must describe the operator's responsibility, and the operator's training and performance should be monitored, quoting, like any other manual process.

11:44 Sam That clause is smarter than it looks. If you lean on the human to lower the testing burden, the human becomes part of the validated system. Same idea as FDA's human-AI team.

11:54 Sarah And the draft's grading starts from a described intended use, not from the model. Section three point one says the intended use and the specific tasks the model assists or automates should be described in detail, including a comprehensive characterisation of the input data and their common and rare variations, which it calls the input sample space, with a process subject matter expert responsible for that description and for the acceptance criteria. So the person who grades the application under Annex Twenty-two is the process expert, not the data scientist.

12:29 Sarah And one more Annex Twenty-two line that grades by comparison rather than by tier: section four point three, no decrease. Acceptance criteria for a model should be at least as high as the performance of the process it replaces, which means you have to know the performance of the manual process first.

12:49 Sam Which most sites don't. My read for a global site: Annex Twenty-two isn't a grading you map onto the others; it's a gate you pass through after them. Grade the model on influence and consequence, then ask two European questions: is the application critical, and is the model static and deterministic. If yes and no, stop.

The AI Act in one takeaway

13:08 Sam The AI Act. Legal's memo. One takeaway, you promised.

13:11 Sarah One. Article Six defines high-risk two ways. First, an AI system that is a safety component of a product, or is itself a product, covered by the Union harmonisation legislation in Annex One and requiring third-party conformity assessment. Medical devices are on that list, so AI inside a regulated device can be high-risk. Second, the use cases in Annex Three, which are eight areas: biometrics, critical infrastructure, education, employment, essential services, law enforcement, migration and border control, and justice. Pharmaceutical manufacturing is not on either list. So almost no manufacturing AI is high-risk under the Act unless it is a safety component of a regulated device.

14:02 Sam And therefore.

14:04 Sarah Therefore the memo is correct and irrelevant to the GMP question. The Act's tiers grade an AI system as a product on the EU market. They say nothing about whether an application is critical under Annex Twenty-two, or what a model's influence on a quality decision is. Not high-risk under the AI Act and should not be used under Annex Twenty-two can both be true of the same agent on the same day.

14:28 Sam That's the whole AI Act segment. Moving on, as promised.

BioPhorum redefines model influence

14:32 Sam BioPhorum. This is the one that reuses FDA's words.

14:35 Sarah BioPhorum's guidance from June twenty twenty-six, written by eight member companies with six more contributing, proposes one harmonised GxP AI risk framework. Section eight point three assesses risk in three elements: scope, meaning system level or individual feature; source and characteristics, meaning where risk comes from and how it shows up; and analysis, meaning consequence and likelihood. Decision consequence is rated low for internal efficiency, moderate for data integrity, process or compliance impact, and high for patient safety, product quality or regulatory impact. So far that is FDA's consequence term with a scale attached.

15:18 Sam And influence?

15:20 Sarah Here it departs. Section eight point five says, quoting, model influence is directly determined by model maturity and serves as a proxy for likelihood. Maturity is two things: autonomy, how independently the system acts, from human in the loop to human on the loop to no oversight; and adaptiveness, from rules-based to static to dynamic. A three-by-three matrix gives the influence rating. Static model with a human in the loop: low. Static with no human oversight: high. Dynamic with no oversight: high. Then influence times consequence gives a composite risk, again three by three.

16:03 Sam So FDA's influence is how much the decision leans on the model. BioPhorum's influence is how likely the model is to go wrong. Same word.

16:12 Sarah Same word, different quantity. FDA and M15 define influence as the weight of the model's evidence relative to other evidence. BioPhorum defines it as the propensity to deviate from intended behaviour, and uses it where an FMEA would put likelihood. They overlap in the easy cases: a model with a real human in the loop rates low under both. They part in others. A rules-based system with no human oversight is moderate influence under BioPhorum's matrix, because it can't drift. Under FDA it is the sole determinant of the decision, so influence is high.

16:50 Sam Which means a number called model influence can't be carried from a BioPhorum assessment into an FDA credibility plan. You re-derive it.

16:58 Sarah Yes. Two more BioPhorum points connect to tonight. It says independent decision-limiting controls, naming independent release testing, orthogonal verification, redundant in-process sampling and process interlocks, may constrain autonomy and so lower the influence rating, provided they operate independently of the AI output. That's the fill-volume logic reaching the industry framework. And section eight point four takes a position against Annex Twenty-two: model type, probabilistic versus deterministic, is deliberately not an independent risk dimension, because those risks are, quoting, already captured through the autonomy and adaptiveness dimensions that define model influence. It calls that position consistent with GAMP and FDA guidance, and says model type should instead raise the validation evidence required.

17:50 Sam So eight companies have put in writing that the European gate is the wrong shape. That's the argument for next episode. Tonight, note that BioPhorum's own matrix still puts a dynamic model with no oversight at high, so the disagreement is about prohibition versus proportionate control, not about the grade.

18:10 Sam Restate. Five gradings on the table. ICH twenty eleven grades the model by role in quality. FDA and M15 grade the decision by influence times consequence. Annex Twenty-two grades the application, critical or not, then gates by model type. The AI Act grades the product and has nothing to say about GMP. BioPhorum grades each feature by consequence times maturity and calls maturity influence. Now the crosswalk that lets a site use one language.

Four tiers for the season

18:42 Sarah The crosswalk defines four tiers on the two factors nearly every framework converges on: influence of the model on the decision, and consequence if the decision is wrong, with Annex Twenty-two's criticality as the GMP anchor. Tier zero, exploratory: the output informs development or business decisions only; no GxP decision, no regulatory evidence. Tier one, advisory in GxP: the output is one input among several; humans or other controls decide; a wrong output is caught downstream.

19:18 Sam And the top two.

19:19 Sarah Tier two, contributing control: the output materially shapes a quality decision but is not the sole determinant; design space models, most in-process controls, feed-forward adjustments, anomaly alerts an operator acts on. Tier three, determinative: the output is the significant or sole indicator of quality, or the sole basis of a GMP decision; a release surrogate, a chemometric assay, autonomous control, batch disposition. Tiers two and three are both critical under Annex Twenty-two; the line between them is ICH's sole-indicator line.

20:02 Sam And the other frameworks land on them how?

20:05 Sarah ICH twenty eleven: low impact at tier zero and one, medium at tier two, high at tier three. FDA and M15 model risk: low, low to medium, medium to high, high. EMA's axes: patient risk rises with the tier when the model touches the process; regulatory impact depends on whether the output reaches a submission.

20:27 Sam And the two that don't grade the model.

20:29 Sarah Annex Twenty-two: non-critical at tiers zero and one, critical at two and three, and at three, if the model is dynamic, probabilistic or a language model, should not be used. The AI Act: minimal risk at every tier unless the model is a device safety component. BioPhorum reads on through consequence: low at tier zero, moderate with a human in the loop at tier one, high with a human in or on the loop at tier two, high with no oversight at tier three. Note that BioPhorum scores at feature level after an AI Act screen, so one system can hold features at several tiers.

21:10 Sam One thing the tiers don't carry. EMA's second axis.

21:13 Sarah Right. Regulatory impact, whether the output becomes evidence in a submission, is absent from the twenty eleven grading and from Annex Twenty-two, and the crosswalk lists that as a gap. A design space model or a stability model can sit at tier one for patient risk and still be high regulatory impact because a reviewer's decision rests on it. The tiers are a GMP scale. When a model's output travels into a dossier, EMA's second question still has to be asked separately.

Four systems, placed and moved

21:44 Sam Good. Now our four systems, placed and then moved. Soft sensor first, as an operator alert.

21:50 Sarah As an alert, the soft sensor predicts a cell density and shows the operator a trend; the operator decides against limits and the offline sample confirms. Tier one. ICH twenty eleven would call it low to medium impact; FDA, low influence, so model risk low to medium; Annex Twenty-two, critical if the step affects product quality, and the draft gives no credit for the low influence; BioPhorum, static model with a human in the loop, low influence, high consequence, composite medium.

22:23 Sam Now make it the release surrogate. The sensor's titre prediction replaces the offline measurement.

22:29 Sarah Tier three. ICH twenty eleven: high impact, exactly its surrogate example, with parallel testing against the reference method through the lifecycle. FDA: sole determinant, influence high, consequence high, model risk high. Annex Twenty-two: critical, and the model must be static and deterministic to be used at all. BioPhorum: static with no human oversight, influence high, consequence high, composite high. Every framework agrees at the top.

23:04 Sam Batch monitoring model. This is ICH's own example, so it should be easy.

23:08 Sarah It is. Used for continuous process verification alongside traditional release testing, tier two: ICH twenty eleven says medium impact in so many words; FDA, medium influence, medium to high consequence, medium risk; Annex Twenty-two, critical, because it touches product quality and data integrity; BioPhorum, static with a human on the loop, moderate influence, composite high. Move it to supporting a real-time release surrogate and it is tier three: ICH says high impact, again in so many words, and everything else follows the soft sensor's path.

23:47 Sam Fill-volume camera.

23:49 Sarah Tier two with the sample test, tier three without, exactly as FDA graded it: medium, then high. ICH twenty eleven would say medium, then high. Annex Twenty-two: critical either way, and a vision classifier is static and deterministic, so in scope and permitted. BioPhorum: the release test is one of its named decision-limiting controls, so it lowers the autonomy rating; remove it and the camera is static with no oversight, high.

The agent's tier

24:20 Sam And the agent. Place it honestly.

24:22 Sarah Drafting a deviation investigation for an engineer who genuinely reviews and re-derives it: tier one, advisory in GxP. ICH twenty eleven has no category for it; it isn't a process model. FDA, from last episode: low influence if the review is real, high consequence, model risk medium. The AI Act: not high-risk. BioPhorum: a pinned foundation model is static; with a human in the loop, influence low; consequence high, because the outcome is a quality decision; composite medium. Four frameworks, roughly one answer.

25:04 Sam And the fifth.

25:05 Sarah Annex Twenty-two. If deviation management is a critical GMP application, and it has direct impact on product quality and data integrity, then a language model should not be used in it, and the tier doesn't matter. If a site argues the drafting step is non-critical because the human decides, the draft allows it with a human in the loop, and the crosswalk names that argument as the open question.

25:29 Sam Now move it.

25:30 Sarah Let it propose corrective actions that get executed without independent re-derivation, or let it clear records. Tier two heading to three. BioPhorum's matrix moves it to human on the loop, influence moderate, composite high; and its governance section says AI systems should not execute critical or regulated actions without a qualified human, and, quoting, must not directly execute electronic signatures. FDA's influence goes high. Annex Twenty-two was already out.

26:03 Sam So for three of our four, the tier is a grade you defend. For the agent, the tier is an argument about criticality that you have to win first, and the way you win it is the way you design the human review. My view: build the agent so that it can't move tiers by accident. No execution, no signature, a review that leaves a trace. Then you're arguing about tier one, which is an argument you can have.

So what for biomanufacturing

26:27 Sam So what for biomanufacturing. Opinions, labelled. First: set the tier before you open a single guidance. The tier tells you which documents apply and how much evidence to gather. A tier one alert needs a different reading list from a tier three surrogate, and most of the wasted validation effort I've watched came from treating everything as tier three because nobody had said which it was.

26:51 Sarah The documents support the ordering. ICH twenty eleven scales validation elements and documentation to the impact grade. Q9 scales formality to uncertainty, importance and complexity. FDA scales the credibility plan to model risk. BioPhorum scales control depth to composite risk, and says which axis drove it should decide which controls you add: monitoring and autonomy constraints if influence drove it, governance and validation evidence if consequence did.

27:24 Sam Second: use the twenty eleven grading as the spine of every validation argument. It's ICH-endorsed, it predates the hype, and its examples are our examples: the batch model at medium with release testing, high as a surrogate. Write the FDA and BioPhorum terms as columns beside it, not instead of it. Third: for a European site, the tier is necessary and not sufficient. After the tier, ask the two Annex Twenty-two questions, critical and model type, and expect the answer to be a gate, not a gradient.

27:57 Sam Fourth: when a number called model influence moves between documents, re-derive it. A BioPhorum influence rating is a likelihood; an FDA influence rating is a weight of evidence. Put both in the assessment and label them. Fifth, for the agent: the grading exercise is the criticality argument. Decide which uses you will defend as non-critical, design the review so the claim stays true, and write down what would move the agent up a tier so that nobody does it by accident. And sixth, Q9's line, which I'd frame: resource constraints are not a reason for less formality. If the tier says high, the tier says high.

28:41 Sarah Who the documents put in the room for the grading: Annex Twenty-two's personnel principle names process subject matter experts, quality, data scientists, IT and consultants, and makes a process expert responsible for the intended use and the acceptance criteria. The Points to Consider put the grade in the regulatory submission, so regulatory CMC signs it too.

Recap and next episode

29:05 Sam Three things. Sarah.

29:07 Sarah ICH graded models in twenty eleven: low, medium, high by their role in assuring quality, with sole indicator as the line and the traditional release test alongside the model as what holds it below that line. Validation and documentation scale with the grade. It is about models generally, which is why it still fits.

29:28 Sam Second, mine: the seven gradings grade different things. The model, the decision, the application, the product, the feature. Impact in M15 isn't impact in the implementation guide, and influence in BioPhorum isn't influence in FDA. Sort by what's being graded before you compare numbers.

29:48 Sarah Third: four tiers for the rest of the season. Exploratory, advisory, contributing control, determinative, defined by influence and consequence, with Annex Twenty-two's critical-or-not as the GMP gate on top. The three classical systems move up the tiers as their orthogonal controls are removed. The agent's tier is an argument about criticality that has to be won first.

30:15 Sam Next time: model type. Can it learn after deployment, and can it be a language model? Why the model you choose is now a compliance decision, and where FDA and the EU disagree. Annex Twenty-two in full, six pages. Are we there yet? We know how much it matters. Now we find out what we're allowed to build.