AI in Biopharma Manufacturing: Are We There Yet?

Transcript — Humans: Who Is Accountable When the Model Is Wrong?

This is the script the AI voices read, so it matches the audio word for word; times are from the render. Researched, scripted and voiced by AI systems under Jack Prior's direction. Sam and Sarah are AI characters; nothing in the episode is Jack speaking, and none of it is a statement of his views or his employer's. Generative AI can be confidently wrong — check the sources. Corrections: jack@jackprior.ai.

00:00 Narrator This is Are We There Yet — a podcast on the evolution of AI in biopharma manufacturing, directed by Jack Prior. A word about how it's made. This episode was researched, scripted and voiced by AI systems. Jack sets the questions and the frames; the AI reads the documents and writes the conversation you're about to hear, between two AI characters: Sam, who plays a manufacturing-science practitioner, and Sarah, who has read the documents. Nothing you hear is Jack speaking, and none of it is a statement of his views or his employer's. Like any generative AI output, it can be wrong — confidently wrong, or missing a nuance — which are exactly the risks this industry is working to mitigate, and exactly what this season is about. Check the sources before you rely on anything. Corrections are welcome at jackprior dot A I. Now, the episode.

00:53 Sam Two forty in the morning, a bioreactor suite, day nine of a run. The soft sensor recommends a feed adjustment. The operator checks the number against the range in the procedure. It's inside. Accept, log, move on. Titre comes in low. The investigation finds that a new lot of a raw material had pushed the sensor's inputs outside anything it was trained on, and the model, doing exactly what models do, extrapolated. The risk assessment for that sensor sits at medium rather than high, and the reason on the form is one line: operator reviews every recommendation.

01:28 Sarah And the investigator's finding?

01:30 Sam The human-in-the-loop control functioned as designed. And I think that sentence is the whole episode. It functioned as designed, and it caught nothing, because it was designed to check something a computer could have checked, by a person who had no way of knowing that the model, not the process, was the thing that was wrong.

01:48 Sarah That's tonight's frame, Sam. Humans. Who is accountable when the model is wrong: what human oversight means across the documents, and why FDA wants the human-AI team evaluated rather than the model alone.

02:03 Sam And the question we've parked twice, since episode four. Which human, with which competence, evaluating which inputs. Let's read.

Which human, and why now

02:12 Sam Sarah, the question and the documents.

02:14 Sarah What does human oversight mean across the documents, and why does FDA want the human-AI team evaluated? Six documents speak to it, from three layers. From law and regulation: the EU AI Act, whose Article Fourteen is titled human oversight, and draft Annex Twenty-two. From regulator guidance: EMA's reflection paper on AI in the medicinal product lifecycle, tonight's anchor, and FDA's January twenty twenty-five draft. From harmonisation on the device side: the Good Machine Learning Practice principles, now an International Medical Device Regulators Forum text. And from industry practice: BioPhorum's June twenty twenty-six risk guidance, the only one of the six that defines its oversight terms.

03:02 Sam Provenance and maturity of the anchor.

03:04 Sarah EMA's reflection paper on the use of artificial intelligence in the medicinal product lifecycle. Issued by the Committee for Medicinal Products for Human Use with its veterinary counterpart; drafted July twenty twenty-three, consulted to the end of that year, adopted final on the ninth of September twenty twenty-four. A reflection paper: the agency's current experience and expectations across the whole lifecycle, and what an EU assessor has read. Its manufacturing section is one paragraph; its oversight content sits in the deployment, governance and ethics sections. It made human oversight an EU expectation across the lifecycle a year before Annex Twenty-two wrote the phrase into GMP.

03:50 Sam Why now, quickly, because I said in episode one why this season exists. The soft sensor and the batch model have had a human looking at them for twenty years, and nobody wrote down what the human was for; the engineer who built the model was usually the one watching it. What changed is the agentic assistant. It drafts the deviation, ranks the root causes, proposes the CAPA, and every document that lets it in the door lets it in on one condition: a human in the loop. The human is now the load-bearing wall of the whole argument for generative AI in GMP. And nobody has specified the wall.

04:33 Sarah Which is where last episode left us. Data can be made to carry the evidence; tonight asks who is standing on it. And episode five's parked question is FDA's: evaluate the human-AI team, but with which human?

04:47 Sam Start with the words, because human oversight turns out to be three different things.

Three meanings of human oversight

04:51 Sarah Three, and most documents don't say which. BioPhorum's June twenty twenty-six guidance defines them in its glossary, citing ISO. Human-in-the-loop, which BioPhorum labels human approval: a person must review, approve, modify or block the output before any consequential action. Human-on-the-loop: the system operates autonomously, and a person supervises its behaviour and can intervene. And a third the glossary calls human supervision, which overlaps the second: a person oversees but does not approve every decision. Its risk matrix uses three columns: human-in-the-loop, human-on-the-loop, and no human oversight.

05:33 Sam And the column you're in changes the grade.

05:35 Sarah Directly. In BioPhorum's framework, autonomy is one of two axes of model maturity, the other being adaptiveness, and maturity is what BioPhorum calls model influence, its likelihood term; we untangled that from FDA's model influence in episode three. A static model with a human in the loop is rated low influence. The same model with a human on the loop is moderate. With no human oversight, high. A dynamic model is moderate even with a human in the loop, and high in the other two columns. So in this framework the oversight setting is worth one full grade.

06:14 Sam Which is precisely why a risk assessment that just says HITL is dangerous. The engineer who reviews the batch model's output the next morning is on the loop, not in it. The batch has moved on. If the form says in-the-loop, the form is claiming a grade it hasn't earned. First rule tonight: when a risk assessment says human-in-the-loop, ask which of the three it means, and whether the human is actually in front of the decision before the decision happens.

06:42 Sarah BioPhorum adds a caution on the same page. Human oversight is one mechanism by which autonomy may be limited, but so are other decision-limiting controls; its examples are independent release testing, orthogonal verification, redundant in-process sampling and automated interlocks, and those should count when assigning autonomy. The human is not the only way to buy down the grade, and, as we'll see, often not the best one.

AI Act Article 14: what an overseer must do

07:10 Sam Now what the documents actually require. Start with the fullest, even though it mostly doesn't reach us.

07:17 Sarah The EU AI Act, Regulation twenty twenty-four sixteen eighty-nine, in force since August twenty twenty-four. Article Fourteen applies only to high-risk AI systems, and, as we established in episode three, almost no pharma manufacturing AI is high-risk unless it is a safety component of a regulated device. The high-risk obligations have also been pushed out: as I understand it, to December twenty twenty-seven for the Annex Three categories, by this summer's Digital Omnibus. So read Article Fourteen for what it is: the most complete statement anywhere of what human oversight is supposed to consist of, from the one document that had to write it as law.

07:59 Sam Read it.

08:00 Sarah Paragraph one: high-risk systems must be designed, including with appropriate human-machine interface tools, so that natural persons can effectively oversee them while in use. Paragraph two: oversight aims to prevent or minimise risks that emerge in intended use or foreseeable misuse, in particular, quoting, where such risks persist despite the application of other requirements. Paragraph three: measures commensurate with the risks, level of autonomy and context of use, built in by the provider or implemented by the deployer.

08:38 Sam Paragraph four is the one.

08:40 Sarah Paragraph four lists what the persons assigned to oversight must be enabled to do. Five things. To properly understand the system's relevant capacities and limitations and to monitor its operation, including detecting and addressing anomalies, dysfunctions and unexpected performance. To remain aware of the possible tendency of automatically relying or over-relying on the output, which the Act names in brackets: automation bias. To correctly interpret the output, taking into account the interpretation tools and methods available. To decide, in any particular situation, not to use the system, or to disregard, override or reverse its output. And to intervene in its operation, or interrupt it through a stop button or a similar procedure.

09:28 Sam Notice what that list is. Understand the limitations, interpret the output, detect unexpected performance. That is not a bounds check. That's a description of somebody who knows how the model behaves. The Act doesn't say who that is, but it's the closest any document tonight comes to describing model expertise as a requirement of the overseer. Hold that thought.

Annex 22: the operator as a manual process

09:49 Sarah Draft Annex Twenty-two, the binding-track GMP text: drafted by EMA's inspectors working group with PIC/S, consulted July to October twenty twenty-five, roughly thirteen hundred comments, no revised text published after the June workshop, though EMA says it is revising the draft; as I understand it, final to the Commission around the end of this year and effective around twenty twenty-seven. It uses the phrase human-in-the-loop in three places, and in the two operating clauses the word it uses for the human is operator.

10:27 Sam Clause one first.

10:29 Sarah Clause one, the scope. Generative AI and large language models should not be used in critical GMP applications. If used in non-critical applications, personnel with adequate qualification and training should always be responsible for ensuring that the outputs from such models are suitable for the intended use, and the draft glosses that phrase as: that is, a human-in-the-loop. So for a generative model outside critical use, the human is not a mitigation. The human is the condition of use.

11:03 Sam Two point one.

11:04 Sarah Two point one, Personnel, names the parties: close cooperation between process subject matter experts, QA, data scientists, IT and consultants, during algorithm selection, training, validation, testing and operation, all with adequate qualifications, defined responsibilities and appropriate level of access. That is the fullest cast list in any of the six documents, and notice who is on it: data scientists, alongside the process experts.

11:35 Sam And three point three, which is where I parked the question in episode four.

11:39 Sarah Three point three, Human-in-the-loop, in the intended-use section. Where a model is used to give an input to a decision made by a human operator, and where the effort to test such model has been diminished, the description of the intended use should include the responsibility of the operator. And then: the training and consistent performance of the operator should be monitored, quoting, like any other manual process. Clause ten point five completes the pair from the operation side: in the same situation, records should be kept, and depending on the criticality of the process and the level of testing of the model, that may imply a consistent review or test of every output, according to a procedure.

12:23 Sam Read the trade that clause makes, because it's honest. You may test the model less because a human is in the loop. In exchange the human becomes a manual process: written responsibility, training, monitored performance, records. The operator as a unit operation with its own qualification. Right instinct. What it doesn't say is what the operator is doing when they review. Which is the question.

12:49 Sarah It doesn't. It says operator, and adequate qualification and training. And between two point one, which names data scientists, and three point three, which names only the operator, the draft never says which of the parties it listed is the one in the loop.

FDA and GMLP: evaluate the human-AI team

13:05 Sam FDA.

13:06 Sarah FDA's January twenty twenty-five draft, still a draft, says nothing about oversight as a requirement and one precise thing about it as evidence. In step four B, the model evaluation part of the credibility assessment plan, among the items to include, quoting: if the context of use involves a human in the loop, ensure that the evaluation methods consider the performance of the human-AI team, rather than just the performance of the model in isolation. That's the sentence we parked in episode five.

13:37 Sam And it's the sentence that forces the question, because you cannot evaluate a team without picking the human. The device principles?

13:45 Sarah The Good Machine Learning Practice principles: written by FDA, Health Canada and the UK's MHRA in twenty twenty-one, adopted by the International Medical Device Regulators Forum as document N eighty-eight in January twenty twenty-five. Principle seven: the device is assessed with a focus on human-AI interactions in the intended use environment, including the performance of the human-AI team, rather than just the device in isolation. And its explanation is the one place in tonight's reading that lists the ingredients: human factors considerations including user skills, user expertise, user understanding of the model outputs and limitations, potential for overreliance, level of device autonomy, and user error, for normal use and reasonably foreseeable misuse.

14:41 Sam User expertise, understanding of the outputs and limitations, overreliance. That's the AI Act's paragraph four in device language. Both describe a person who knows the model, and both stop short of saying so.

14:55 Sarah And step seven, briefly: when credibility falls short of the model risk, FDA's listed outcomes include downgrading the model influence with additional evidence, and establishing controls to mitigate risk. A human in the loop is one such control. Whether it earns the downgrade is what the evaluation of the team is meant to show.

15:17 Sam Checkpoint, a third of the way in. Three meanings of oversight, and the column moves the grade. The AI Act's five capacities, which read like model expertise without saying so. Annex Twenty-two makes the operator a manual process but doesn't say what they're checking. FDA and the device principles say evaluate the team, and the device text names expertise and overreliance. Now the anchor.

EMA: a risk plan for every model

15:44 Sarah The reflection paper comes at oversight from the other end: it asks what happens when there isn't any. Section two point five point six, model deployment: for all models, especially those where there is no human-in-the-loop, a system risk management plan should be developed that defines likely risks of failure modes of the algorithm. That plan covers the consequences of incorrect predictions or classifications, the monitoring and mitigation or correction approaches, how to trigger a suspension or decommissioning of the model, and how to carry it out.

16:19 Sam So EMA's default is: assume the model will fail, write down how, and write down how you'd stop it. The human is an especially, not the plan.

16:27 Sarah Yes. Two point two puts responsibility for ensuring that all algorithms, models, datasets and data processing pipelines are fit for purpose with the sponsor, applicant, marketing authorisation holder or manufacturer. Two point six, governance: standard operating procedures implementing GxP principles on data and algorithm governance should be extended to all data, models and algorithms used for AI where regulatory impact or patient risk is high. And two point eight lists seven ethics principles borrowed from the Commission's high-level expert group; the first is human agency and oversight, and the section says a human-centric approach should guide all development and deployment.

17:12 Sam The manufacturing paragraph.

17:13 Sarah Two point three point six, one paragraph. AI in manufacturing, from process design and scale-up to in-process control and batch release, is expected to increase; model development, performance assessment and lifecycle management should follow quality risk management principles, with patient safety, data integrity and product quality in view; consider ICH Q eight, Q nine and Q ten, awaiting revision of current regulatory requirements and GMP standards. That last clause is a pointer to Annex Twenty-two before it existed.

17:50 Sam One more from it, Sarah, because I think it's the most useful sentence in the paper for tonight, and it's in the conclusion.

17:56 Sarah The conclusion says efforts should be made in all organisations to, quoting, reciprocally integrate data science competence with the respective fields within medicines development, manufacturing and pharmacovigilance. Reciprocally. Data science competence into manufacturing, and manufacturing competence into data science. It is the only sentence in the six documents that treats the competence question as the organisation's to solve.

18:23 Sam And a precedent a few pages earlier that I'd steal. For AI in precision medicine, the paper asks for guidance that prescribers can critically apprehend, and fall-back treatment strategies in cases of technical failure. Translate that to a plant: the operator must be able to critically apprehend the recommendation, and there must be a fall-back when the model fails. Critically apprehend is the competence question in six syllables.

The oversight paradox

18:51 Sarah Now the paradox, and the documents state it themselves. BioPhorum's governance section: human oversight plays a dual role in AI risk management and can function as both a risk control and a risk factor. Oversight mechanisms such as human-in-the-loop are essential, it says, but they may also introduce variability, bias or informal workarounds if roles, competencies and decision boundaries are not clearly defined. And on adaptive models, citing Annex Twenty-two: reliance on human intervention alone is insufficient as a long-term control, because system behaviour may evolve outside the assumptions of the original validation state.

19:33 Sam And the sentence after that is the one I'd put on the wall.

19:36 Sarah Effective AI governance must explicitly define when and how human oversight functions as a validated control, and when it constitutes an additional source of risk requiring training, procedural controls and ongoing monitoring. The AI Act's contribution is the name for the failure mode, automation bias, and the requirement that overseers remain aware of it. And Annex Twenty-two's answer, from three point three, is to treat the human as a manual process with its own monitored performance.

20:09 Sam Put those three together and you get a rule that isn't written in any of them and follows from all of them: a control you cannot measure is not a control. If the operator's review is a control, it has a performance: a rate at which it catches model errors, measured against something. If nobody has measured it, and nobody can say what it would catch, it isn't a control. It's a signature.

20:34 Sarah BioPhorum's own worked example makes the point. Its appendix grades a quality-system deviation module feature by feature. For the ranked root-cause suggestions it lists overreliance by users as a source of risk, and mandatory investigator justification as a control. For the proposed corrective and preventive action it rates consequence high and influence moderate, with human approval in place, and the composite risk still comes out high. The document's own illustration that the loop does not automatically buy the grade.

Which human, which competence, which inputs

21:09 Sam So here's the question, properly this time. It's mine to put, Sarah, and you tell me where the documents land. A human-in-the-loop claim asserts that a person will catch what the model gets wrong. What inputs is that person evaluating that could not have been coded into the system?

21:26 Sarah Split it.

21:27 Sam Two cases. Case one: the review means check that the output, or the diagnostics, are within defined bounds. The recommendation is in range; the confidence is above threshold; the inputs are inside the training envelope. Every one of those is a rule, and a rule is something a computer enforces better than a person, every time, at three in the morning. Case two: the review means stop something unexpected or inappropriate. That's a judgment, and it needs one of two competences. Process expertise: this looks physically wrong for this reactor, this product, this day; something an experienced operator may genuinely have more of than the data scientist. Or model expertise: the model is misbehaving, extrapolating, attributing to the wrong features. And model expertise does not exist on the night shift of a twenty-four-seven operation.

22:18 Sarah Take case one against Annex Twenty-two, because the draft has already given the bounds checks to the system. Three point one: the intended use should include a comprehensive characterisation of the input sample space, and a process subject matter expert is responsible for it. Nine point two: the model should have a threshold, and if the confidence score is very low it should flag the outcome as undecided rather than make an unreliable prediction. Ten point three: system performance monitored against its metrics. Ten point four: monitor whether the input data are still within the model sample space and intended use, with metrics defined for drift. The draft assigns the envelope, the confidence and the drift to the system. It does not leave them for the operator.

23:11 Sam So if the operator's job is case one, the operator is doing the system's job, worse than the system would, and adding paperwork, not control.

23:20 Sarah And for case two, which competence: none of the six says. Annex Twenty-two says operator, and personnel with adequate qualification and training. BioPhorum says a qualified human. FDA says the human-AI team. The AI Act's paragraph four, correctly interpret the output and understand the capacities and limitations, and the device principles' user expertise and understanding of the model outputs and limitations, are the closest anyone comes to naming model expertise. And the reflection paper's reciprocal integration is the closest anyone comes to saying whose job it is to make that expertise exist.

24:01 Sam Four questions, then, for every HITL line on every risk assessment. Which human. Which competence. Evaluating which inputs. Measured how. If the answers are the operator, qualified, the screen, and we don't, the line is decorative.

Accountability does not move

24:15 Sarah Before your answer, Sam, the part that doesn't change. None of these documents moves accountability. The quality unit's responsibilities under the US regulations, part two hundred eleven, and under EU GMP chapter one are not in tonight's reading, and no AI text we've read alters them — FDA's draft says in a footnote that the quality control unit's responsibilities under two eleven point twenty-two remain. FDA's draft says in a footnote that sponsors remain responsible for compliance with statutory and regulatory requirements regardless of the technology utilised. EMA, as we heard, places it with the sponsor, applicant, marketing authorisation holder or manufacturer. BioPhorum adds two things: an AI owner, identified during development, who stays accountable for fitness for use throughout the lifecycle; and AI systems must not directly execute electronic signatures, and should not be permitted to execute critical or regulated actions without intervention from a qualified human.

25:22 Sam So the model never signs. A person signs. And the practical question, the one the investigator should have asked at the cold open, is: who signed, on the strength of what, and could that person have known the model was wrong. If the honest answer to the third is no, the risk assessment relied on a control that could not function, and the accountability sits with whoever approved that assessment, not with the operator who accepted a number that was in range.

Sam's position: encode, name, escalate

25:48 Sam My position, then. It's the character's synthesis, offered for you to test, and the documents supply every premise and none of the conclusion. Three moves. One: encode every check that can be encoded. Bounds, input-distribution checks, out-of-distribution detection, confidence thresholds, drift metrics. Annex Twenty-two already asks for most of them. The human is never the bounds checker.

26:14 Sarah Two?

26:15 Sam Two: write into the intended use, which three point three already requires, the specific judgment the human supplies, by competence rather than by job title. Not the operator reviews. Rather: a person able to recognise physically implausible behaviour of this process reviews; or, a person able to recognise model misbehaviour reviews. Three: if the judgment you've written down is model expertise, and the plant does not staff it at the hours the decision is made, the control is fiction, and the risk grade must not get credit for it. Either build an escalation path to model expertise, a named competence on call and an undecided state that holds the decision until they answer, or accept the higher tier and the evidence burden that comes with it.

27:05 Sarah What the documents support in that: Annex Twenty-two's three point three requires the responsibility in the intended use and the operator's performance monitored; BioPhorum requires governance to define when oversight is a validated control; the AI Act's paragraph four describes the capacities; FDA requires the team to be evaluated. What none of them says is that a control that cannot be staffed earns no credit. That's the step you're adding.

27:35 Sam It's the step a Q nine risk assessment would take for any other control. A pressure interlock that isn't wired doesn't reduce a risk score. A person who can't tell that the model is wrong isn't wired.

Four systems, three questions

27:46 Sam Second checkpoint, then the four systems. So far: three meanings; what each document requires; the paradox; the four questions; accountability stays put; and encode, name, escalate. Now the old friends, each with the human-in-the-loop claim its risk assessment probably makes, and three questions apiece: which human, which competence, which inputs.

28:10 Sarah The soft sensor, with an operator reviewing recommendations. Which inputs? If the operator compares the recommendation to a range, that is Annex Twenty-two's three point one and nine point two, and the system should enforce it. If the operator is asked whether the recommendation makes physical sense for this run, that is process expertise, and it is real: an experienced operator can know that a feed increase at this hour of this phase is wrong before any model can. That judgment belongs in the intended use, and three point three says monitor it like a manual process: measure how often the operator correctly overrides.

28:46 Sam And it's measurable. Seed the review with known-bad recommendations and see what gets caught. You'd do that for a visual inspector; do it here. The competence is process, the plant staffs it on every shift, and the credit is earned. That's the good case.

29:02 Sarah The multivariate batch model, reviewed by an engineer the next morning. That is human-on-the-loop in BioPhorum's terms, not in-the-loop: the batch has progressed. The engineer's competence may well be model expertise, the right person, but the review comes after the decision. On BioPhorum's matrix a static model with a human on the loop is moderate influence, not low.

29:28 Sam Name it honestly and grade accordingly. On-the-loop is fine; most monitoring works that way. What's not fine is claiming the in-the-loop grade for it. And if you want the in-the-loop grade, the model has to be able to hold the batch, which means the undecided state and somebody on call.

29:45 Sarah The fill-volume vision system, with an operator confirming rejects. The failure mode is the one the AI Act names: automation bias. After a thousand agreements the operator agrees with the thousand-and-first. Under three point three the operator is now a manual inspection process, and plants have qualified manual inspectors for decades: known defects in the stream, a catch rate, requalification. The competence is process. The measurement already exists.

30:16 Sam Two things. FDA's own example already had the better control: the sample fill-volume test at release, which is orthogonal and doesn't get tired; BioPhorum's decision-limiting controls list says the same. The operator confirming rejects buys less risk than the sample test does. And if the operator can't actually override the reject, they aren't a control at all; paragraph four says the overseer must be able to disregard, override or reverse.

30:47 Sarah And the agentic assistant: a drafted deviation investigation, ranked root causes, a proposed CAPA, approved by QA. Under Annex Twenty-two clause one it is non-critical use only, with the human as the condition. Which inputs is the reviewer evaluating? Factual accuracy and completeness against the source records: did the batch record say what the draft says it says; did the draft miss the alarm at hour fourteen; is the proposed action the one the evidence supports. Which competence? Reading the records is QA's own. Judging whether the model has invented a plausible root cause is closer to model expertise, and the reflection paper's warning about generative models producing plausible but erroneous or incomplete output, written for product information, applies word for word.

31:40 Sam And the inputs question has a physical answer: does the reviewer have the source records open, or only the draft? If only the draft, the reviewer is checking whether the story is coherent, which is exactly what a language model is best at producing. The control is then measuring fluency. BioPhorum's example asks for citation requirements and back-testing against closed deviations, and that's the right shape: every claim in the draft points at a record, and the reviewer checks the pointer, not the prose.

32:10 Sarah The signature stays human; BioPhorum says AI must not execute electronic signatures, and nothing in the regulations lets it. And the night-shift test, Sam.

32:20 Sam Who, at three in the morning, can tell that the model rather than the process is wrong? For the soft sensor: the operator can tell the process is wrong, the system can tell the inputs have drifted, and nobody on shift can tell the model is misbehaving; so the intended use says the model's own checks carry that, and where they can't, the decision waits. For the agent: nobody on shift can tell an invented root cause from a real one without the records; so the agent doesn't propose actions on the night shift, or it proposes them into a queue a competent reviewer opens at seven. That's not a limitation of the technology. It's a limitation of the staffing, and writing it down is the control.

So what for biomanufacturing

33:01 Sam So what for biomanufacturing. Opinionated, and labelled as mine. Rewrite every HITL line in your risk assessments as four fields: which human, which competence, which inputs, measured how. Then sort. Where the answer is a bounds check, automate it and stop claiming credit; the Annex will make the system do it anyway. Where the answer is process expertise, keep it, write it into the intended use, and monitor it like any manual process: seeded errors, a catch rate, requalification. Where the answer is model expertise you don't staff at the hours the decision is made, the credit goes away and the tier goes up, or you build the escalation path: an undecided state that holds the decision, and a named competence on call.

33:48 Sarah Who has to be in the room, from the documents: Annex Twenty-two's two point one list, process experts, QA, data scientists and IT, plus BioPhorum's AI owner. And the place the human gets defined is the intended-use document; three point three puts it there, and it is the same document FDA calls the context of use.

34:11 Sam Three consequences. MSAT and process science own the process-expertise judgments, and can turn them into a qualification: what does a competent reviewer notice, and how do we know they noticed it. QA owns the signature, so the question could I have known is QA's, and the agent's reviewer needs the records open and a procedure that checks pointers, not prose. Regulatory CMC owns FDA's evaluate-the-team: the credibility plan has to name the human and the test, and if you validate the team with the data scientist and deploy it with the operator, you validated a team you don't run.

34:48 Sarah And the data point that runs under every episode: the model's own checks, the drift metrics, the input envelope, the undecided flag, are computed from the data the last episode was about. If the basement isn't dry, the system can't catch what you've just taken away from the operator.

35:06 Sam Which is the thesis, and it's mine: the human is a validated control or it is not a control. The documents give it every premise and none of the conclusion.

Recap and next time

35:15 Sam Three things. Sarah.

35:16 Sarah First, human oversight means three different things, in-the-loop, on-the-loop and none, and BioPhorum shows the setting is worth a full grade of model influence. When a risk assessment says HITL, ask which one it means, and whether the person is in front of the decision.

35:34 Sarah Second, the documents converge on the same shape without naming it: the AI Act's five capacities, the device principles' user expertise and overreliance, Annex Twenty-two's operator monitored like a manual process, FDA's human-AI team. None says which competence. The reflection paper says building it is the organisation's job, reciprocally.

35:58 Sam Third, mine: encode every check a computer can run, name the judgment the human supplies by competence, and if that competence isn't on the night shift, the control is fiction and the grade must not get credit for it. Escalate, or accept the tier. And accountability never moved: the signature is human, and could that person have known is the question.

36:22 Sam Next time: change. What happens when the model, or the process, changes? Post-approval change management protocols, predetermined change control plans, Established Conditions, and the vendor update nobody planned for; the foundation model under the agent that got swapped on a Tuesday. Are we there yet? We know who signs. Next we ask what happens when what they signed for changes underneath them.