Beyond Manual Literature Screening: From AI Automation to Decision Preparation in Pharmacovigilance in Europe

Andreas Hofmann, Managing Director, p.AI GmbH

Artificial intelligence can increasingly extract, structure, and compare pharmacovigilance-relevant information from scientific literature. The key question is no longer whether AI can support literature assessment, but how responsibilities should be divided between AI and pharmacovigilance professionals.

Introduction:

As of August 2026, the EU AI Act is generally applicable, although some obligations for high-risk AI systems apply at later dates. For pharmacovigilance organisations, a key consideration is to determine the role and classification of each AI system under the Act, rather than assuming that AI used in PV is automatically classified as high-risk. AI-literacy requirements are already applicable, while provider and deployer obligations should be assessed based on the system's actual architecture, intended purpose, and use case. 

AI functions are appropriately validated, and meaningful human oversight is built into the workflow.

Literature assessment is more than screening

Scientific literature remains an important source of safety information in pharmacovigilance. A single publication may contain an Individual Case Safety Report (ICSR), information relevant to an existing safety concern, or findings that contribute to understanding a medicinal product's safety profile.

However, literature assessment is rarely as simple as reading an abstract and deciding whether it is "relevant" or "irrelevant." A publication may discuss several medicinal products, adverse events, patient populations, and varying levels of information about exposure, outcomes, and causal relationships. Decision preparation must classify study types (human clinical reports vs. non-clinical/bench research) upfront. Filtering out non-clinical studies before entity extraction prevents unnecessary ICSR evaluation pipelines. Study type should be classified upfront, distinguishing human clinical reports from non-clinical or bench research. 

It is also important to distinguish literature monitoring as an end-to-end process from assessment performed after identifying a potentially relevant publication.

Search strategies, database surveillance, and article retrieval answer the question:

What should be reviewed?

Literature assessment answers:

What does publication mean for the safety of a company's medicinal products?

Much of the work after retrieval involves collecting, structuring, standardising, and comparing information that a pharmacovigilance expert needs before making a decision. This creates a significant opportunity for AI-assisted decision preparation.

AI should automate the preparation of safety decisions, not accountability for them.

Literature assessment is not one task

The term "literature assessment" can hide several distinct activities.

First, relevant information must be identified, including publication type, medicinal products, adverse events, indications, dose, route of administration, patient characteristics, exposure, and outcomes.

Second, relationships need to be understood. Identifying both amoxicillin and rash in a publication is not enough. The reviewer needs to determine whether the rash was associated with amoxicillin, another medicine, a drug interaction, or an unrelated part of the publication. 

Third, information often needs to be standardised and placed into context. Free-text adverse events may need to be mapped to controlled terminology such as MedDRA. Medicinal products or substances must be connected with the relevant portfolio, while findings may need comparison with applicable reference safety information.

Only then do more consequential questions arise:

  • Does the publication contain a potential ICSR?
  • Is the event serious?
  • Is the relationship considered causal?
  • Is the finding already reflected in applicable safety information?
  • Could it affect the known safety profile?
  • Does it require further action?

Where AI can remove repetitive work

The strongest near-term applications of AI are tasks where the expected output is relatively well-defined and can be verified against the source document.

Modern language-processing systems can classify publications and extract structured information traditionally requiring manual reading and transcription. Relevant fields include medicinal products, adverse events, dose and exposure, patient characteristics, indications, outcomes, and authors' safety conclusions.

For literature assessment, the extraction should prioritise data elements needed to determine whether a publication contains a valid ICSR. At minimum, workflow should identify:

  • Identifiable patient
  • Identifiable reporter/primary source
  • Suspect or interacting medicinal product
  • Suspected adverse event/adverse drug reaction (AE/ADR)

These four elements form core minimum criteria for an ICSR under ICH E2D(R1) and GVP Module VI. Assessment should also capture seriousness criteria where applicable, along with relevant clinical details needed for case evaluation. Authors’ safety conclusions may provide useful context, but they should not take precedence over these regulatory-critical ICSR elements.

The next step is relationship extraction.

Simply identifying entities has limited value if the system cannot preserve their context. A useful workflow should distinguish between a publication that merely mentions a medicinal product and an adverse event and one that actually reports a relationship between them.
AI can also support terminology standardisation by suggesting appropriate MedDRA concepts for extracted events. This can reduce repetitive work, although human oversight remains important. 

MedDRA Term Selection: Points to consider guidance recognises the need for human oversight when IT tools perform term selection, ensuring that the resulting term reflects reported information and makes medical sense.

The objective should therefore not be autonomous medical assessment:

AI identifies and structures evidence.

Controlled terminology standardises it.

The PV specialist verifies the interpretation.

This model can reduce mechanical information handling without transferring medical accountability to an algorithm.

From extraction to contextual assessment

Information extraction is only the first layer of useful decision support.
Consider an AI system that identifies:

Medicinal product: Substance A
Adverse event: Gastrointestinal haemorrhage

The extraction may be accurate, but it does not answer questions that matter to a pharmacovigilance professional.

Is Substance A part of a relevant product portfolio? What role did it have in the reported case? Which product information applies? Does the terminology used in the publication correspond to the terminology in the reference document?

The value of AI, therefore, increases when extraction is connected to context.

An advanced workflow can move from an unstructured publication to a drug-event relationship, normalise information against controlled terminology, connect it with relevant products, and compare findings with applicable reference safety information.

AI's role is not to make regulatory determinations. Its role is to ensure that relevant information, terminology, comparisons, and source passages are already assembled when the expert begins the assessment.

A risk-based Human–AI workflow

The future of literature assessment is unlikely to be a choice between humans and AI. A more useful question is:

Which responsibility should sit where?

Structured, low-ambiguity tasks that can be readily verified may support a higher degree of automation. Examples include metadata extraction, publication classification, and identification of source passages.

Activities involving greater medical or regulatory consequences should retain stronger expert control. These include ambiguous seriousness assessments, causality assessments, interpretations of changes to safety profiles, benefit-risk implications, and regulatory action.

A practical allocation looks like this:

This is a design principle rather than a regulatory classification:

As ambiguity, medical judgement, and consequence increase, human control should increase.

This direction is consistent with CIOMS Working Group XIV concept of "intelligence augmentation", which combines human and artificial intelligence to reduce repetitive work while retaining expert involvement in activities such as differential diagnosis and causality assessment.

Validation should follow the task

A risk-based workflow also changes how AI performance should be evaluated.

Saying that an AI system is "95% accurate" provides little useful information without specifying the task, dataset, and consequence of an error.

For literature triage, sensitivity may be particularly important. A system that reduces workload but systematically misses relevant safety publications has failed where it matters most.

For entity extraction, appropriate measures may include precision, recall, and field-level accuracy. For drug-event relationship extraction, identifying both entities is insufficient; the system must preserve their correct relationship.

For terminology mapping, relevant measures could include whether the correct candidate is ranked first and how frequently experts override suggestions.

Reference-document comparison should also be decomposed into individual failure points. Errors may result from event extraction, terminology normalisation, selection of applicable reference information, or comparison itself.

The validation question should therefore be:

Does this specific AI-supported function perform sufficiently well for its intended role within a controlled workflow?

Not simply:

Does the model appear intelligent?

Operational metrics such as override rates, assessment time, escalations, publication-type performance, and trends support ongoing monitoring. Applicants/MAHs should ensure algorithms, models, datasets, and processing pipelines are fit for purpose and compliant with applicable legal, GxP, scientific, and technical standards. This is with risk-based fit-for-purpose validation/performance assessment within the validated process or system and lifecycle monitoring. 

Human oversight must be designed into the system

"Human in the loop" is often presented as a solution to AI risk, but the phrase alone provides little assurance.

Meaningful human oversight requires reviewers to have the information and controls necessary to challenge the system. A pharmacovigilance professional should be able to inspect source evidence, see the AI proposal, understand uncertainty, accept or reject the result, and modify it where necessary.

Reviewer corrections can also reveal real-world failure modes. An AI system may perform well on straightforward publications but struggle with complex reviews, tables, specific document types, or particular terminology.

This feedback can support monitoring, controlled improvement, and future validation. It should not imply uncontrolled continuous learning in production; changes to models, terminology, or processing logic require appropriate governance.

A reviewer who receives only a generated conclusion and an Approve button is not exercising meaningful oversight.

Human oversight is therefore not just a model-governance issue. It is also a workflow and interface-design issue.

The system should make verification easier than blind acceptance.

From automation to decision preparation

The potential of AI in pharmacovigilance literature assessment is not simply that it can read documents faster.

Its greater value may be in changing what reaches the expert.

Instead of beginning with an unstructured publication, pharmacovigilance professionals can increasingly begin with organised evidence: relevant medicinal products, adverse events, relationships, terminology, reference-information comparisons, confidence indicators, and source passages already assembled for verification.

This does not remove professional judgement. It changes where professional time is spent.
The success of AI-assisted literature assessment should therefore not be measured only by the number of manual steps eliminated. More meaningful questions are:

  • Does relevant evidence reach reviewers faster?
  • Are proposed conclusions traceable to source evidence?
  • Is inconsistent information handling reduced?
  • Is expert attention concentrated on cases where human judgement provides the greatest value?

The likely next stage of literature assessment is neither fully manual nor fully autonomous.

AI prepares evidence. Experts verify/review the interpretation and retain responsibility for the decision.

The goal is not autonomous pharmacovigilance. It is a system in which experts spend less time assembling evidence and more time applying their expertise to it. 
Transitioning from full automation hype to structured decision preparation allows safety leaders to manage expanding literature volumes while maintaining robust, audit-ready human oversight. 

References

  1. International Council for Harmonisation. ICH E2D(R1): Post-Approval Safety Data – Definitions and Standards for Management and Reporting of Individual Case Safety Reports. Step 4, 2025.
  2. European Medicines Agency. Reflection paper on use of Artificial Intelligence (AI) in medicinal product lifecycle. EMA/CHMP/CVMP/83833/2023. Final version, September 2024.
  3. MedDRA Maintenance and Support Services Organisation. MedDRA Term Selection: Points to Consider. Release 4.26, March 2026.
  4. Hakim JB, Painter JL, Ramcharran D, et al. need for guardrails with large language models in pharmacovigilance and other medical safety critical settings. Scientific Reports. 2025.
  5. Council for International Organisations of Medical Sciences. Artificial Intelligence in Pharmacovigilance. CIOMS Working Group XIV Report. December 2025.
Andreas Hofmann

Andreas Hofmann is Managing Director of p.AI GmbH, a Germany-based company developing AI-assisted software for pharmacovigilance. His work focuses on applying artificial intelligence to regulated drug-safety workflows, particularly literature assessment, structured safety-data processing, traceability, and human–AI decision support.