Why AI Is Changing What Good Validation Looks Like
Jason Bryant, General Manager of AI Platforms, ArisGlobal
As AI becomes embedded in everyday pharmacovigilance, validation is only the starting point. This article examines how CIOMS XIV reframes responsible AI around continuous monitoring, proportionate oversight and lifecycle governance, and considers what pharmaceutical organisations must do to ensure AI systems remain safe, effective and appropriately governed throughout routine operation.
The comprehensive landmark report and framework set out seven principles for using AI safely in pharmacovigilance, together advocating a risk-based approach to governance, validation and more. Here, ArisGlobal’s Jason Bryant highlights some of the guidance’s main assumptions, as recently distilled in a podcast with Denny Lorenz, a core member of the CIOMS XIV working group.
When CIOMS Working Group XIV finalised and formally published its guidance on trusted AI in pharmacovigilance at the end of last year[1], this was the culmination of an international consensus process that brought together regulators, academic researchers and the life sciences industry. Its central premise is that an AI system cannot simply earn trust once and keep that badge by default. Rather, trust has to be regularly renewed as long as the system is in service. The guidance expresses this through seven linked principles: a risk-based approach, human oversight, validity and robustness, transparency, data privacy, fairness and equity, and governance and accountability.
What a risk-based approach requires
The framing of risk underpins every other principle in the guidance, for good reason. Where an indiscriminate and potentially heavy-handed approach to checks could risk the promised gains (better use of resources, contained costs), a risk-based approach to AI-based system and process governance, validation and so on involves weighing how likely a problem is against how serious the consequences would be if one arose. It then uses these risk assessments to identify, prioritise and manage anything that could affect how a PV system behaves. Key considerations include whether the AI’s output feeds a high-stakes decision, and whether a human checks that output before it’s acted on (vs the system running largely unsupervised). Focusing effort where it matters lets organisations use AI properly without unnecessary risk. But this also needs reassessing over time, e.g., whenever a system’s performance changes enough to warrant it.
Misconceptions about human oversight
Human oversight is one of the harder principles to get right. CIOMS makes a careful distinction. Oversight can mean human-in-the-loop, where a person and the system jointly produce a decision, or it can mean human-on-the-loop, where the system decides and a person checks the result afterwards. But neither will be effective without defining acceptable performance in advance or testing the system against realistic data. None of this happens automatically just because a reviewer exists. A drug safety team could assign a reviewer to every single case and still be missing a documented PV system master file, still have gaps in data privacy, and still have no way of showing whether that reviewer is catching anything that matters. Oversight only becomes genuine evidence for a risk-based approach once the reviewer’s own performance is itself being tracked.
Validation beyond go-live
Validity and robustness is where CIOMS is most prescriptive, and where the guidance most directly challenges the way that validation has worked traditionally. Performance needs to be demonstrated under realistic conditions, the guidance states using a properly representative range of data, patients and event types. This matters in PV, since AI is frequently asked to catch rare events or unusual patterns - circumstances where a narrow test set is most likely to miss a real weakness. Provisions here need to extend well beyond deployment, since a model’s inputs vary from case to case, and a system that seems unchanged on the surface might actually be running an updated prompt or a new underlying version.
Transparency, privacy, fairness and governance
According to the CIOMS guidance, transparency means disclosing enough about a system, its design, and its performance that someone outside the project could understand a result well enough to question it. In life sciences, especially when it comes to patient reports, data privacy is treated as close to sacred. Since large language models carry a real risk of re-identifying patients from supposedly anonymised data, that risk has to be assessed before rather than after deployment. Fairness and equity, meanwhile, mean checking that training and test data sufficiently represent the populations a medicine will actually reach (skewed reference data is a common cause of biased AI output). Proper governance and accountability require clearly-defined roles so that if something goes wrong, someone specific is answerable for it.
Putting the guidance into practice
For those deciding whether a use case is ready for production, CIOMS includes a governance grid - a structured set of questions to probe whether risk has been formally assessed, whether an oversight process actually exists on paper, and to what extent governance is ongoing. This isn’t about holding back progress, as long as any shortfalls are known and somebody assumes responsibility for fixing them.
As to when a PV subject matter expert should be brought into an AI project, Working Group discussions concluded in favour of early involvement, on the basis that identifying a model’s blind spots before a new use case for AI has been honed beats finding them once the tool is already running.
Similar logic has a bearing on how manual review is expected to evolve. Currently, even though a single patient adverse event report can carry two or three hundred distinct fields, most organisations still check each one, irrespective of how reliably the AI extraction has already performed. A viable interpretation of CIOMS’ guidance here (it suggests that the frequency, amount, or depth of human controls may be gradually reduced as confidence in routine AI performance increases) might be to treat twelve consecutive months of dependable results in the field as a reasonable threshold for scaling checks back to just the low-confidence extractions. Although regulators haven’t confirmed that they will accept a narrower approach, it is something companies need to get to grips with. CIOMS frames this as an ongoing conversation to be had with regulators.
The same direction of travel is in evidence beyond PV parameters, with the FDA and EMA jointly publishing their own guiding principles - covering the full medicines lifecycle at the start of 2026[2]. Both plot validation and oversight against risk, and expect monitoring to run for a system’s entire working life, tracking closely with what CIOMS laid out for pharmacovigilance a month earlier.
Whichever recommendations companies consult, the important point is that none of the various principles work in isolation, nor can they be satisfied once and then forgotten. Trust in an AI system must be re-earned with every new model version, every update, every additional year a system stays in production. Going forward, this is the standard against which pharmacovigilance will be measured against.
References:
[1]Council for International Organizations of Medical Sciences (CIOMS), ‘Artificial Intelligence in Pharmacovigilance’, CIOMS Working Group XIV report, Geneva, December 2025. Available at:https://cioms.ch/working_groups/working-group-xiv-artificial-intelligence-in-pharmacovigilance/
2European Medicines Agency and U.S. Food and Drug Administration, ‘EMA and FDA set common principles for AI in medicine development’, 14 January 2026. Available at: https://www.ema.europa.eu/en/news/ema-fda-set-common-principles-ai-medicine-development-0