ICH M14 Explained: Real-World Data for Safety Assessment

Explains the ICH M14 framework for generating regulatory-grade real-world evidence from non-interventional studies, including research questions, data fitness-for-use, feasibility, design, bias, analysis, reporting and the relationship with EU PASS requirements.

Take test

ICH M14 Explained: Real-World Data for Safety Assessment

Real-world data are often described as if their value comes from scale: millions of patients, years of healthcare activity and the possibility of studying medicine use outside controlled trials. ICH M14 takes a more disciplined position. A large dataset is useful only when its data are sufficiently relevant and reliable for the specific safety question, the study design can address that question credibly, and the remaining limitations are understood well enough to support regulatory interpretation.

That is the central idea of ICH M14.

The guideline establishes harmonised principles for planning, designing, analysing and reporting non-interventional studies that use real-world data for medicine safety assessment. It does not turn every database analysis into regulatory evidence, and it does not replace regional pharmacovigilance legislation or the EU procedures governing post-authorisation safety studies. Instead, it provides a scientific framework for deciding whether a proposed real-world-data study can generate evidence adequate for the regulatory question being asked.

In the EU, ICH M14 became legally effective as Step 5 guidance on 18 March 2026. EMA has also stated that GVP modules affected by ICH M14 will be revised and that, in the interim, the ICH guidance should be applied where it affects existing GVP guidance.

Purpose and Scope

ICH M14 focuses on non-interventional pharmacoepidemiological studies that utilise real-world data for the post-marketing safety assessment of medicines.

The guideline is intended to improve consistency across regions by providing common principles for:

The objective is not simply methodological neatness. Harmonisation can reduce the need for several regulators to request different studies addressing the same concern and can make one well-designed study more usable across jurisdictions.

ICH M14 uses the term should to describe recommendations rather than legal requirements. That distinction matters. In the EU, some non-interventional post-authorisation safety studies are subject to binding legal or procedural requirements under pharmacovigilance legislation and GVP Module VIII. M14 supplies scientific and methodological guidance; it does not itself determine whether a study is an imposed PASS, what submission pathway applies, or which statutory timelines govern it.

Real-World Data and Real-World Evidence

Real-world data (RWD) are data relating to patient health status or healthcare delivery that are generated in routine settings rather than solely through conventional randomised clinical trials.

Examples can include:

Real-world evidence (RWE) is the evidence generated through analysis of real-world data.

The distinction is fundamental:

Real-world data
      ↓
Study design + definitions + analysis
      ↓
Real-world evidence
      ↓
Regulatory interpretation

A database is not evidence by itself. Evidence arises only after the data are used within a study designed to address a defined question.

What ICH M14 Is Trying to Solve

Non-interventional studies have long been used in pharmacovigilance. The problem is that regulatory confidence can be undermined when different studies use incompatible definitions, poorly justified data sources, inadequate control of bias or analyses that were adapted after results became known.

M14 addresses this by making the reasoning chain explicit:

Safety concern
      ↓
Research question
      ↓
Key design elements
      ↓
Initial feasibility
      ↓
Draft protocol
      ↓
Detailed data fitness assessment
      ↓
Is the study adequate?
   ┌──────┴──────┐
   │             │
  Yes            No
   │             │
Finalise      Alternative data /
protocol      primary collection /
   │          hybrid design
Execute
   ↓
Assess limitations
   ↓
Final report

Although this sequence looks linear, M14 emphasises that study development is iterative. Feasibility findings can change the question, design or data source. Early regulatory interaction can also alter the proposed approach.

The important principle is that the research question should drive the study and data selection, not the other way around.

The EU Regulatory Position

ICH adopted M14 at Step 4 on 4 September 2025. EMA subsequently adopted the Step 5 guideline with a legal effective date of 18 March 2026.

EMA currently lists M14 within the pharmacovigilance ICH framework and has stated that GVP modules affected by M14 will be revised. Until that revision is complete, M14 should be applied as far as it affects GVP.

The relationship can be understood as follows:

Framework Primary role
EU pharmacovigilance legislation Binding legal obligations
GVP Module VIII EU guidance for PASS classification, conduct, submission and oversight
ICH M14 Harmonised methodological principles for non-interventional studies using RWD
ENCePP methodological guidance Additional scientific detail and best practice
ICH E2D(R1) ICSR and post-approval safety-data reporting implications during studies

This is why an EU study may need to satisfy several frameworks simultaneously.

A non-interventional PASS may have a protocol that follows GVP procedural expectations while also applying M14 principles for data fitness, design and analysis. If the study generates individual safety information, ICH E2D(R1) and the applicable regional ICSR rules may also become relevant.

What Is in Scope

M14 applies primarily to RWE submitted to regulators for post-marketing medicine safety assessment.

The guideline can accommodate studies that use:

Primary data collection can be important where routine data do not contain a necessary variable, outcome definition, confounder or special-population characteristic.

This does not convert the study into an interventional trial merely because additional information is collected. The defining issue remains whether treatment assignment is determined by the protocol.

What Is Outside the Main Scope

M14 explicitly excludes several categories from its principal scope.

These include:

The guideline also recognises that artificial intelligence, pharmacogenomics and other evolving technologies may be relevant to RWD and RWE, but it does not provide comprehensive guidance for those technologies.

This boundary prevents M14 from being treated as a universal standard for every activity involving “real-world” information.

Descriptive and Inferential Objectives

A study should make clear whether its objective is descriptive or causal/inferential.

A descriptive study may estimate:

An inferential study asks whether exposure is associated with an outcome in a way that supports a causal interpretation.

That distinction changes the design burden.

A causal question normally requires a scientifically justified comparator, control of confounding, careful construction of time at risk and an explicit interpretation of remaining bias. A descriptive study may not require the same causal machinery, but it still requires a valid population, clear definitions and fit-for-use data.

The next step is therefore not “choose a database”. It is formulate the research question precisely enough that the minimum data requirements become visible.

Start With the Research Question

M14 treats the research question as the anchor for the entire study.

A useful research question identifies, as relevant:

M14 notes that structured frameworks such as PICOTS can help formulate the question. The purpose is not to force every study into a template. It is to make the intended comparison explicit enough that the study design can be evaluated before data analysis begins.

For an inferential question, the researcher should also be clear about the causal contrast being estimated.

For example:

Among adults with atrial fibrillation initiating Medicine A or Medicine B, what is the comparative incidence of major bleeding during the first 180 days of treatment?

That question immediately creates data requirements. The study needs to identify new treatment initiation, distinguish the two medicines, recognise atrial fibrillation, define major bleeding, establish follow-up, determine treatment changes and measure important confounders.

By contrast:

Is Medicine A associated with bleeding?

is too broad to define a defensible study.

Use Existing Knowledge Before Designing the Study

M14 recommends critical appraisal of existing information before finalising the research question.

This can identify:

This step reduces unnecessary duplication and prevents a new study from reproducing a known design flaw.

For a safety concern already examined in several claims-database studies, for example, the main gap may not be sample size. The unresolved problem may be incomplete laboratory data, poor disease-severity measurement or inability to confirm the clinical outcome.

The new study should therefore be designed around the remaining uncertainty, not simply around the availability of another large database.

Feasibility Is a Scientific Assessment

One of the most important contributions of M14 is that feasibility is treated as part of scientific design rather than as a procurement exercise.

A feasibility assessment asks whether candidate data can support the minimum elements needed to answer the question with acceptable validity and precision.

M14 describes at least two phases:

  1. an initial scan to identify potentially suitable data sources and narrow the options;
  2. a more detailed assessment of the candidate sources.

The detailed assessment should establish whether the needed variables are actually present, operationally usable and sufficiently valid. M14 also makes an important separation between feasibility and outcome analysis: feasibility can use information such as patient counts, overall outcome frequency, covariate availability and expected statistical precision, but it should not compare outcomes between exposed and comparator groups in a way that reveals the study result before the protocol is finalised.

This means that feasibility is not equivalent to asking:

Does this database contain ten million patients?

A database with ten million patients may still be unusable if exposure cannot be reconstructed accurately, important confounders are absent or the outcome is poorly captured.

Define Minimum Data Requirements Before Choosing the Data Source

A disciplined approach begins by stating what the study needs.

Typical minimum requirements include:

Population

Exposure

Comparator

Outcome

Covariates

The research question should therefore generate the data specification.

Fit-for-Use: The Core M14 Concept

M14 uses fit-for-use for the determination that a proposed data source is sufficiently relevant and reliable for a particular study.

This is intentionally question-specific.

A database is not globally “fit for use”. It may be highly suitable for one study and inadequate for another.

For example, a national dispensing database may be excellent for studying treatment initiation and persistence but weak for a safety outcome that requires detailed imaging findings.

A hospital EHR may capture laboratory-defined acute kidney injury well but may not represent outpatient medicine exposure reliably.

Fit-for-use therefore has two principal dimensions:

Dimension Core question
Data relevance Does the data source contain the people and variables needed for this question?
Data reliability Can those data be trusted sufficiently for the intended analysis?

Data Relevance

M14 links relevance to matters such as:

Relevance is not simply completeness.

A variable can be present for every patient yet still be irrelevant to the clinical construct the study needs.

Data Reliability

M14 identifies reliability concepts including:

Data provenance concerns where the data originated and how they were recorded, captured and added to the source.

This matters because identical-looking fields can have very different meanings depending on how they were generated.

A diagnosis code may represent:

Without understanding provenance, the researcher can misinterpret the variable even when the dataset is technically complete.

Data Recency and Refresh

Data fitness also depends on timing.

M14 recommends considering:

These issues matter when a safety question is urgent.

A highly detailed database that becomes available only after a 12-month lag may be less useful for an emerging safety concern than a somewhat less detailed source refreshed weekly.

This creates a legitimate design trade-off between:

The trade-off should be explicit rather than hidden.

Understand How the Data Were Generated

Routine healthcare data are produced for purposes other than the research study.

Examples include:

The data therefore reflect the processes that created them.

Claims data may capture reimbursed services well but lack clinical nuance.

EHR data may contain rich clinical information but be fragmented when patients receive care across different systems.

Registries may contain disease-specific detail but represent selected patients or centres.

The researcher should therefore understand:

Healthcare process
      ↓
Data generation
      ↓
Coding / transformation
      ↓
Research dataset
      ↓
Study variable

Error can enter at every stage.

Common RWD Source Types

Electronic Health Records

EHRs can provide diagnoses, clinical notes, laboratory values, vital signs, procedures and prescribed medicines.

Potential limitations include:

Administrative Claims

Claims can be powerful for longitudinal healthcare utilisation and dispensing patterns.

Strengths can include large populations, structured coding and relatively complete capture of reimbursed services.

Limitations can include:

Registries

Registries can contain clinically rich disease-specific data and long-term follow-up.

Their suitability depends on:

Secondary use of registry data requires the same fit-for-use reasoning as other secondary healthcare data.

Digital Health Technologies

Digital health technologies can generate measurements that are difficult to obtain through conventional databases.

M14 notes that these data should undergo the same fit-for-use assessment as other sources.

Relevant considerations can include:

The mere fact that a measurement is digital or continuous does not make it more reliable.

Research Networks and Federated Data

Large safety studies increasingly use networks in which data remain at participating institutions while standardised queries are executed across nodes.

Examples of federated or distributed networks include large regulatory and academic infrastructures.

The advantage is that researchers can analyse several populations without transferring all patient-level data to one central location.

The methodological challenge is consistency.

The study needs sufficient harmonisation of:

A common data model can standardise structure without eliminating differences in underlying clinical practice or data provenance.

When Existing RWD Are Not Enough

M14 explicitly allows the feasibility process to conclude that the available RWD are inadequate.

That is a successful scientific outcome, not a failure.

If a required confounder or outcome cannot be measured adequately, alternatives can include:

A useful operational principle is:

Do not weaken the safety question merely to fit the database.

Instead, determine whether the question can be answered with acceptable validity and, if not, redesign the evidence-generation strategy.

From Feasibility to Protocol

Once the research question and candidate data sources are sufficiently understood, M14 moves into protocol development.

The protocol should make the study's reasoning traceable. It should explain not only what will be done but why the chosen population, exposure, comparator, outcome, covariates, design and analysis are appropriate for the safety question.

That rationale is important because many observational-study decisions are defensible only in context.

For example, a 30-day risk window may be reasonable for an acute arrhythmic effect but inappropriate for a malignancy hypothesis. The number “30” is therefore not meaningful without the biological and clinical reasoning behind it.

Choose the Design That Fits the Question

M14 recognises several common non-interventional designs, including:

No design is universally preferred.

Cohort Designs

A cohort study follows exposed and comparison groups over time and can estimate incidence or rates directly.

It is often useful when:

Case-Control Designs

A case-control study begins with patients who experienced the outcome and compares prior exposure with controls.

It can be efficient for rare outcomes but requires careful attention to:

Self-Controlled Designs

Self-controlled approaches compare risk periods within the same individual.

They can control automatically for fixed patient characteristics but rely on assumptions about:

The design should therefore be justified by the clinical and temporal characteristics of the safety concern.

Target Population and Study Population

M14 distinguishes the population about which the researcher wants to draw conclusions from the population actually represented in the data.

The target population is the population relevant to the regulatory question.

The study population is the population that can actually be constructed from the selected data source and eligibility criteria.

The difference matters for external validity.

A safety question may concern all older adults receiving a medicine, while the available database may represent only commercially insured patients receiving care in a particular network.

The study can still be informative, but the limitation should be recognised explicitly.

The protocol should therefore address:

Exposure Definition

Exposure is rarely as simple as “medicine present”.

Depending on the source, exposure may be inferred from:

Each represents something different.

A prescription records intent to treat.

A dispensing record shows that medicine was supplied.

An administration record may provide stronger evidence that treatment was actually given.

The protocol should define:

These choices should reflect the suspected biology.

For an acute adverse reaction, the relevant period may begin immediately after treatment. For an outcome with long latency, cumulative or historical exposure may matter more.

New Users and Prevalent Users

Whether a study includes new users or prevalent users changes interpretation.

New-user designs can align baseline covariate assessment more clearly with treatment initiation and avoid some forms of survivor bias.

Prevalent-user designs may provide larger populations but can preferentially include patients who already tolerated treatment.

Neither choice is automatically wrong. The protocol should explain the consequences.

Comparator Selection

The comparator defines the counterfactual question.

Possible comparators include:

A useful comparator should help reduce differences unrelated to the medicine itself.

For example, comparing a medicine used for severe disease with healthy untreated people can create profound confounding by indication and disease severity.

An alternative medicine used at the same treatment stage may improve comparability, although residual differences can remain.

The protocol should explain why the comparator was selected and which treatment-selection differences are expected.

Outcome Definition

The outcome should represent the clinical event the study is intended to evaluate.

Potential identification methods include:

A code-based algorithm may be efficient but can misclassify cases.

For example, a diagnosis code for liver injury may represent:

The study should therefore consider whether the operational definition has adequate validity for the intended inference.

Covariates

Covariates can include:

The purpose is not to adjust for every available field.

Covariates should be selected because they relate to:

Automated inclusion of hundreds of variables can be useful in some methods, but it does not replace clinical reasoning about the major sources of confounding.

Bias Is a Design Problem, Not a Statistical Afterthought

M14 gives substantial attention to bias and confounding.

This is important because observational studies do not become valid merely by applying a sophisticated model at the end.

Potential problems should be identified during design.

Selection Bias

Selection bias can occur when inclusion, observation or follow-up differs according to exposure and outcome risk.

Examples include:

Information Bias

Information bias occurs when exposure, outcome or covariates are measured inaccurately or differently across groups.

Examples include:

Incorrect handling of time can create strong but artificial associations.

Examples include:

The design should establish the index date, exposure windows, baseline period and follow-up rules coherently.

Confounding

Confounding occurs when treatment groups differ in factors that also affect outcome risk.

Important examples include:

Restriction, matching, stratification, regression, propensity-score methods and weighting can reduce measured confounding.

They cannot guarantee removal of unmeasured confounding.

This is why data fitness and design are inseparable.

A dataset that lacks a critical confounder may remain inadequate no matter how sophisticated the analysis.

Quantitative Bias Analysis

M14 gives quantitative bias analysis a notable role.

The purpose is to estimate how systematic error could affect the direction or magnitude of the observed association.

Potential targets include:

Quantitative bias analysis can be used during design to assess whether a planned study is likely to be informative and later during interpretation to evaluate the robustness of the result.

This does not mean every M14 study must perform every form of quantitative bias analysis.

The important principle is that material sources of uncertainty should be characterised explicitly rather than hidden behind a single adjusted estimate.

Validation of Key Variables

Where the validity of exposure, outcome or other key variables is uncertain, validation may be needed.

Validation can use:

The protocol should explain whether an existing validation study is sufficiently applicable to:

A positive predictive value reported ten years earlier in a different healthcare system is not automatically transferable.

If several candidate definitions are available, M14 encourages evaluation of their performance and the potential effect of misclassification on the study.

Missing Data Begin With Understanding Why Data Are Missing

Missingness is not simply an empty cell.

A variable may be missing because:

The mechanism of missingness can itself create bias.

The protocol should therefore describe:

The next stage is to ensure that the study's data-management and analysis processes preserve the scientific logic defined in the protocol.

Data Management and Curation

M14 treats data management as part of study validity.

A non-interventional safety study should have a data-management or data-curation plan before study initiation. The plan should describe how the analytic dataset will be created, transformed, checked and secured.

This is especially important for RWD because the source data often pass through several layers before analysis:

Clinical / administrative source
        ↓
Data holder
        ↓
Extraction
        ↓
Transformation / common data model
        ↓
Linkage
        ↓
Analytic dataset
        ↓
Statistical analysis

Every transformation creates an opportunity for error.

The study documentation should therefore make it possible to reconstruct:

Data Quality, QA and QC

M14 distinguishes the broader scientific question of whether data are fit-for-use from the operational controls used to preserve data quality during the study.

Potential quality risks include:

QA and QC should be proportionate to the risks that could materially affect the evidence.

This does not mean every RWD study requires the same validation regime as a clinical trial.

The relevant question is:

Which errors could change the study population, exposure, outcome, confounder measurement or final estimate, and how are those errors prevented or detected?

For a distributed network, for example, the highest-value controls may include:

Roles of Data Holders and Researchers

RWD studies often depend on data holders who are not the study sponsor.

A data holder may control:

The researcher may therefore receive a dataset without having direct access to the original clinical record.

M14 places importance on understanding these upstream processes because data reliability depends partly on them.

A scientifically credible study should be able to answer:

Outsourcing data preparation does not transfer scientific responsibility for the study.

Statistical Analysis Should Follow the Research Question

The statistical analysis plan should be developed before outcome analyses are conducted and should align with the research question and protocol.

The objective is not to select the model that produces the strongest association.

The analysis should estimate the quantity relevant to the safety question while making the assumptions explicit.

Potential elements include:

The choice of measure should follow the design and clinical question.

A relative measure alone may not communicate the public-health importance of a safety concern. Where possible, absolute measures can help show how much additional risk is associated with exposure.

Confounding Adjustment

Common analytical approaches include:

These approaches rely on measured data.

If a clinically important confounder is absent, a more complex model cannot create the missing information.

A strong M14 analysis therefore connects confounding control back to the feasibility assessment:

Potential confounder identified
        ↓
Is it measured adequately?
     ┌──────┴──────┐
     │             │
    Yes            No
     │             │
Adjustment     Residual-bias
possible       strategy / alternative
               data source / redesign

Time-Varying Exposure and Confounding

Medicine use changes over time.

Patients may:

A study that classifies exposure only once at baseline may therefore misrepresent the actual risk periods.

Similarly, some covariates change over time and can be affected by earlier treatment.

These situations require careful design and may require methods that account for time-varying exposure or time-dependent confounding.

The correct approach depends on the causal structure and research question. The key M14 principle is that timing assumptions should be explicit rather than embedded invisibly in code.

Missing Data Analysis

The analysis plan should describe the expected missing-data mechanism and how missingness will be characterised.

Potential strategies include:

No method makes missing data harmless automatically.

For example, multiple imputation relies on assumptions about the relationship between observed and unobserved values. If a clinically important variable is systematically missing because it is measured only in the sickest patients, the assumptions require careful evaluation.

The extent and consequence of missing data should therefore appear in the interpretation, not only in a technical appendix.

Sensitivity Analyses

M14 gives sensitivity analyses a central role in evaluating uncertainty.

A sensitivity analysis should test a material assumption.

Examples include changing:

The rationale should be prespecified where possible.

A long list of alternative models is not automatically informative. The useful sensitivity analysis is the one that addresses a plausible source of bias or uncertainty.

Quantitative Bias Analysis

Quantitative bias analysis can estimate how strong a bias would need to be to materially change the interpretation.

For example, if an outcome algorithm has imperfect sensitivity and specificity, the researcher can examine how plausible misclassification affects the estimated association.

Similarly, the study can evaluate the potential impact of an unmeasured confounder under different assumptions.

These analyses are particularly valuable when the central limitation cannot be removed from the data.

Pre-Specified and Post-Hoc Analyses

M14 emphasises the distinction between analyses specified before examining comparative outcomes and analyses added afterward.

Pre-specified analyses should be documented in the protocol or statistical analysis plan.

Post-hoc analyses can be scientifically useful—for example, to understand an unexpected result—but should be clearly identified and justified.

This protects against a common interpretive problem:

Many analyses performed
        ↓
Only favourable / striking result highlighted
        ↓
Apparent certainty greater than actual certainty

Transparency about analytical chronology is therefore part of scientific credibility.

Machine Learning and Derived Algorithms

M14 does not provide comprehensive AI guidance, but it does address machine-learning or other derived methods used within the analysis.

Where such methods are used, the statistical analysis plan should describe relevant matters such as:

This is important because an algorithm can introduce hidden data dependence.

For example, an outcome model trained in one hospital system may perform poorly in another population.

The use of machine learning therefore does not remove the need for validation, provenance and study-specific fitness assessment.

Multiple Data Sources

Combining data sources can improve representativeness or supply missing variables, but it also creates additional complexity.

Differences can occur in:

Pooling data without understanding these differences can produce a precise but misleading average.

Possible approaches include:

The protocol should explain why the chosen approach is appropriate.

Generalisability and External Validity

A statistically valid result in the study population may not apply equally to the target population.

Potential limitations include:

M14 therefore treats representativeness as part of data relevance.

The study report should explain which patients are represented and where extrapolation becomes uncertain.

Special Populations

M14 specifically discusses populations that are often under-represented in pre-authorisation trials, including:

RWD can be especially valuable in these populations, but methodological challenges can be greater.

Examples include:

These studies may require linkage to complementary sources such as pregnancy registries, birth records or disease registries.

The same M14 principles still apply: define the question, identify the minimum information needed, assess data fitness and design around the known sources of bias.

Interpretation Should Integrate Precision and Validity

A narrow confidence interval does not prove that the estimate is correct.

A study can be highly precise and still biased.

Conversely, an imprecise estimate may still be important if the outcome is serious and the data exclude some clinically relevant possibilities.

Interpretation should therefore consider together:

This is why M14's framework evaluates not merely the statistical result but the adequacy of the evidence for the regulatory question.

Reporting, Submission and Regulatory Interpretation

A well-designed study can still lose credibility if the final report hides changes, omits limitations or presents exploratory analyses as though they were prespecified.

M14 therefore extends the same traceability expected during study design into reporting.

The final study documentation should allow a regulator to understand:

The report should connect the numerical result back to the regulatory safety question.

A result is not complete merely because a hazard ratio or risk difference has been calculated.

Adverse Events and Individual Case Reporting During M14 Studies

M14 does not create a separate ICSR reporting system.

It explicitly directs researchers to the applicable post-approval safety-reporting framework and regional requirements.

This creates an important connection with ICH E2D(R1).

The ICSR implications depend partly on how the study data are collected.

For example:

The QPPV.com article ICH E2D(R1): What Changed in Post-Approval Pharmacovigilance explains that distinction in detail.

The M14 protocol should describe how AEs, ADRs, other observations and product-quality complaints will be identified and managed according to the applicable jurisdiction.

Study Documents for Regulatory Submission

Depending on the jurisdiction and study status, regulators may expect documents such as:

M14 encourages early interaction with regulators when the planned study raises important questions about data fitness, design or interpretation.

This is particularly valuable when:

Early dialogue can prevent a technically sophisticated study from later being judged incapable of answering the regulatory question.

Dissemination and Transparency

The study should distinguish between the regulatory submission process and broader dissemination of scientific findings.

Transparency supports reproducibility and trust.

Useful controls include:

Selective publication of only favourable or statistically significant results undermines the value of the evidence base.

Study Documentation and Record Retention

M14 recommends that key study records be handled so that the study can be:

The documentation system should support:

Retention periods are determined by applicable jurisdictional requirements rather than by one universal M14 retention period.

This is another example of the law/guidance distinction: M14 describes the documentation principles, while regional requirements determine the legally applicable retention obligations.

Relationship With EU PASS Requirements

M14 and GVP Module VIII overlap, but they answer different questions.

GVP Module VIII

GVP Module VIII addresses the EU pharmacovigilance framework for post-authorisation safety studies, including:

ICH M14

M14 focuses on whether a non-interventional study using RWD is scientifically capable of generating reliable evidence for the safety question.

A practical distinction is:

GVP Module VIII
"What regulatory PASS framework applies?"

ICH M14
"Can this RWD study generate adequate evidence?"

A study may therefore satisfy its PASS procedural requirements yet still be scientifically weak if the data are not fit-for-use or the design does not control major bias.

Conversely, a methodologically strong database study may not be an imposed PASS and therefore may not follow every procedural pathway applicable to an imposed EU study.

QPPV.com covers the EU classification framework separately in GVP Module VIII: Non-Interventional PASS Requirements.

Relationship With Pharmacoepidemiology

M14 should also be distinguished from general pharmacoepidemiology.

Pharmacoepidemiology provides the broader scientific methods used to study medicine use and effects in populations.

M14 narrows those methods into a harmonised regulatory framework for non-interventional studies using RWD for medicine safety assessment.

The broader methodological foundation is covered in Pharmacoepidemiology in Pharmacovigilance.

This layered architecture prevents duplication:

Pharmacoepidemiology
        ↓
General scientific methods

ICH M14
        ↓
Harmonised RWD safety-study framework

EU GVP Module VIII
        ↓
Regional PASS regulatory framework

ENCePP and the Living Methodological Guide

The ENCePP Guide on Methodological Standards in Pharmacoepidemiology remains an important complementary resource.

In 2026, ENCePP moved the Guide toward a living-resource model, with individual chapters updated as methodology evolves. The last full revision was the July 2023 revision, while chapter-level updates now continue independently.

This is useful because M14 is intentionally high-level.

For detailed methodological questions—such as specific bias-control methods, exposure validation, distributed-network design or pregnancy-study methodology—the ENCePP Guide can provide additional depth.

It should be used as methodological guidance rather than treated as a source of new legal obligations.

Practical M14 Implementation for a Pharmacovigilance Organisation

A mature implementation can be organised around eight controls.

1. Define the Regulatory Safety Question

Start with the signal, safety concern or uncertainty.

State whether the objective is:

Identify what decision the study is intended to support.

2. Translate the Question Into Minimum Data Requirements

Define what must be measured for:

Do this before selecting the database.

3. Perform Structured Feasibility

Use an initial scan to identify candidate sources.

Then perform detailed assessment of:

4. Document Why the Data Are Fit-for-Use

The protocol should explain both:

Statements such as “large validated database” are insufficient unless the validation is relevant to the specific study variables and question.

5. Design Around Bias

Before analysis, identify the main threats to validity.

For each major threat, state whether it will be addressed through:

6. Preserve Analytical Traceability

Maintain:

Distinguish prespecified analyses from exploratory work.

7. Connect Study Results Back to Pharmacovigilance

The final study should not end with a statistical estimate.

The pharmacovigilance interpretation should explain whether the evidence:

8. Integrate Regulatory and QPPV Oversight

Where the study forms part of the pharmacovigilance system, governance should ensure visibility of:

QPPV oversight should be proportionate to the study's significance and its role in the safety system. It does not require the QPPV to perform epidemiological analysis personally.

Worked Example: Evaluating a Thromboembolic Signal

Assume spontaneous reports suggest a possible increase in venous thromboembolism after Medicine X.

Step 1 — Research Question

Among adults initiating Medicine X, is the incidence of venous thromboembolism higher than among patients initiating a clinically appropriate alternative treatment during the first six months?

Step 2 — Minimum Data Requirements

The study needs:

Step 3 — Feasibility

A large claims database contains both medicines and enough events but lacks reliable body mass index.

A linked claims-EHR source contains BMI but covers fewer patients.

The question is therefore not simply which database is larger.

The researchers should assess whether the missing BMI information in the larger source creates material residual confounding and whether linkage or sensitivity analysis can address it.

Step 4 — Design

A new-user cohort design with an active clinical comparator may reduce differences in treatment indication.

The risk window should be clinically justified.

Patients with recent VTE may need exclusion or separate analysis depending on the research question.

Step 5 — Bias Assessment

Important threats include:

Step 6 — Analysis

The primary analysis may use adjusted comparative incidence with predefined sensitivity analyses for:

Step 7 — Interpretation

Suppose the primary estimate suggests increased risk but the association substantially attenuates when obesity is captured in the linked dataset.

The regulatory conclusion should reflect that residual confounding may explain at least part of the association.

The study should not be summarised merely as “positive” or “negative”.

That is the practical value of M14: it structures the reasoning needed to decide how much confidence the evidence deserves.

Inspection and Assessment Perspective

An inspector or regulatory assessor could reasonably ask:

The focus is not whether one preferred epidemiological technique was used.

The focus is whether the evidence chain is scientifically defensible and traceable.

Illustrative Failure Modes

The following are hypothetical examples, not reported inspection findings.

Database First, Question Second

A company buys access to a large database and then searches for a safety question the data can answer.

Why it fails: the data source drives the research rather than the regulatory uncertainty.

“Millions of Patients” Used as Proof of Fitness

The protocol cites database size but does not assess outcome validity or key confounders.

Why it fails: precision cannot compensate for systematic bias.

Comparator Chosen for Convenience

A comparator group is available in the database but differs profoundly in disease severity.

Why it fails: confounding by indication can dominate the association.

Outcome Algorithm Never Validated

A rare serious outcome is identified using codes whose positive predictive value is unknown.

Why it fails: the study cannot quantify how much apparent risk may arise from misclassification.

Protocol Written After Feasibility Results Leak Outcome Differences

Comparative results are explored during feasibility and then the primary analysis is selected.

Why it fails: feasibility becomes outcome-driven study optimisation.

Sensitivity Analyses Used as Decoration

Dozens of alternative models are presented without linking them to identified uncertainties.

Why it fails: volume of analysis is not equivalent to robustness.

Negative Study Interpreted as No Risk

The study has wide confidence intervals and poor outcome capture but is described as disproving the signal.

Why it fails: absence of statistical significance is not proof of absence of a clinically important effect.

Practical Review Checklist

Question What good evidence looks like
Is the safety question precise? Explicit population, exposure, comparator, outcome, timing and objective
Was data selection question-driven? Minimum requirements defined before source selection
Are data fit-for-use? Relevance and reliability assessed for this study
Is provenance understood? Source-generation and transformation pathway documented
Is the population representative? Target vs study population differences explained
Are exposure and outcome valid? Operational definitions justified and validated where needed
Are major confounders measured? Clinically important confounders identified before analysis
Are time-related biases addressed? Index date, baseline, risk windows and follow-up coherent
Is missingness understood? Mechanism, extent and consequences described
Are sensitivity analyses purposeful? Each addresses a material assumption
Are post-hoc analyses transparent? Clearly identified and justified
Is the analysis reproducible? Data transformations, code and software documented
Are limitations integrated into conclusions? Regulatory interpretation reflects residual uncertainty
Is PV action explicit? Result linked back to signal, risk or benefit-risk question

Key Takeaways

References

  1. International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use. ICH M14: General Principles on Planning, Designing, Analysing, and Reporting of Non-interventional Studies That Utilise Real-World Data for Safety Assessment of Medicines. Final Step 4 guideline, adopted 4 September 2025.
    https://database.ich.org/sites/default/files/ICH_M14_Step4_Final_Guideline_2025_0905.pdf

  2. European Medicines Agency. ICH M14 guideline on general principles on planning, designing, analysing, and reporting of non-interventional studies that utilise real-world data for safety assessment of medicines — Step 5. EMA/CHMP/ICH/155061/2024. Legal effective date: 18 March 2026.
    https://www.ema.europa.eu/en/ich-m14-guideline-general-principles-plan-design-analysis-pharmacoepidemiological-studies-utilize-real-world-data-safety-assessment-medicines-scientific-guideline

  3. European Medicines Agency. Good pharmacovigilance practices (GVP). Current GVP framework and planned revisions related to ICH M14.
    https://www.ema.europa.eu/en/human-regulatory-overview/post-authorisation/pharmacovigilance-post-authorisation/good-pharmacovigilance-practices-gvp

  4. European Medicines Agency. Guideline on good pharmacovigilance practices (GVP) Module VIII — Post-authorisation safety studies (Rev. 3). EMA/813938/2011. Legal effective date: 13 October 2017.

  5. European Network of Centres for Pharmacoepidemiology and Pharmacovigilance. ENCePP Guide on Methodological Standards in Pharmacoepidemiology. Living methodological resource; last full revision July 2023 with chapter-level updates thereafter.
    https://encepp.europa.eu/encepp-toolkit/methodological-guide_en

  6. International Council for Harmonisation. ICH E2D(R1): Post-Approval Safety Data: Definitions and Standards for Management and Reporting of Individual Case Safety Reports. Final Step 4 guideline, adopted September 2025.

  7. European Medicines Agency and Heads of Medicines Agencies. HMA-EMA Catalogues of real-world data sources and studies. Current EU catalogue infrastructure for data sources and studies.

Regulatory Note

This article explains ICH M14 as reviewed on 1 October 2026.

ICH M14 provides harmonised scientific and methodological recommendations. It does not replace EU legislation, GVP Module VIII, national requirements, ethical or data-protection obligations, or study-specific regulatory commitments.

The guideline itself specifies that the word should means suggested or recommended rather than required. Where a study is an imposed or otherwise regulated PASS, binding obligations arise from the applicable legal and regional framework.

EMA has stated that GVP modules affected by ICH M14 will be revised. Until those revisions are complete, organisations should apply M14 where it affects GVP while continuing to follow the current applicable EU pharmacovigilance framework.

The worked example and failure modes in this article are illustrative and are not presented as actual regulatory inspection findings.

Revision History

Last reviewed: 2026-10-01

QPPV.com