Disproportionality Analysis in Pharmacovigilance
Spontaneous-report databases contain many thousands or millions of individual case safety reports (ICSRs). At that scale, it is impossible to rely only on manual reading to identify every potentially important product-event pattern. Disproportionality analysis provides a statistical way of focusing attention: it asks whether a particular adverse event is reported with a particular medicinal product more often than would be expected from the reporting pattern in the rest of the database.
The method is powerful because it turns a very large reporting system into a screening instrument. Its limitation is equally important: the underlying data were not collected as a controlled epidemiological study. The number of reports depends on exposure, awareness, publicity, reporting culture, indication, geography, time since launch and many other influences. A high disproportionality statistic therefore means disproportionate reporting, not necessarily increased incidence and not necessarily causation.
That distinction is central to regulatory signal management. GVP Module IX Addendum I describes disproportionality methods as tools for identifying groups of ICSRs that may warrant further investigation. When predefined algorithmic conditions are met, the result is a signal of disproportionate reporting (SDR). The SDR is then reviewed clinically and scientifically to decide whether it represents information that merits signal validation or further assessment.
- Disproportionality Analysis in Pharmacovigilance
- Purpose and Regulatory Context
- The 2 × 2 Reporting Table
- Bayesian and Shrinkage Approaches
- Worked Example and Interpretation
- Data Preparation and Analytical Choices
- Bias, Confounding and Masking
- Thresholds and Monitoring Frequency
- Governance and Practical Implementation
- Potential Failure Modes
- Inspection Considerations
- Practical Checklist
- Key Takeaways
- References
- Regulatory Note
Purpose and Regulatory Context
Disproportionality analysis belongs primarily to signal detection. It is designed to reduce a large search space by ranking or flagging combinations for human review. It does not replace signal validation, prioritisation or assessment.
Within the EU, GVP Module IX provides the overall framework for signal management, while Module IX Addendum I addresses methodological aspects of signal detection from spontaneous reports. EMA's EudraVigilance Data Analysis System (EVDAS) supports signal detection and evaluation and includes a measure of disproportionality based on the reporting odds ratio (ROR).
The regulatory framework deliberately allows methodological flexibility. Addendum I discusses choices such as the disproportionality statistic, thresholds, minimum case counts, subgrouping and monitoring frequency. It does not impose one universal company-wide threshold that must be used for every medicinal product and database. Instead, the signal detection algorithm should be justified for the data and purpose, and its performance should be understood.
From report to statistical flag
A simplified sequence is:
ICSR database → product-event counts → disproportionality statistic → signal detection algorithm → SDR → clinical review → signal validation/assessment.
The distinction between the statistic and the algorithm is important. A statistic such as the ROR is a numerical summary. A signal detection algorithm (SDA) defines the rules by which that statistic, confidence bounds, minimum counts or other criteria are converted into an SDR.
EMA's methodological guidance notes that the choice of thresholds can substantially affect signal-detection performance. Thresholds that are too low produce many false-positive SDRs and consume review resources; thresholds that are too high can delay or miss true adverse reactions. Threshold design is therefore a performance question, not merely a compliance box.
The 2 × 2 Reporting Table
Most frequentist disproportionality measures can be understood from a 2 × 2 table:
| Event of interest | All other events | |
|---|---|---|
| Product of interest | a | b |
| All other products | c | d |
The cells represent report counts rather than exposed patients. The table asks whether the proportion of reports containing the event is higher for the product of interest than for the comparator background.
Reporting Odds Ratio
The Reporting Odds Ratio is:
ROR = (a / b) ÷ (c / d) = ad / bc
An ROR above 1 means that the reporting odds for the event are higher for the product than for the comparator reports. Because estimates based on small counts are unstable, interpretation commonly uses a confidence interval as well as the point estimate.
EVDAS uses the ROR as its disproportionality measure. The existence of an elevated ROR does not itself establish a regulatory signal. The associated cases, medical plausibility, known product information, confounding and other evidence still require evaluation.
Proportional Reporting Ratio
The Proportional Reporting Ratio compares the proportion of reports for the event among reports for the product with the corresponding proportion among other products:
PRR = [a / (a + b)] ÷ [c / (c + d)]
ROR and PRR are mathematically related and often rank combinations similarly when events are uncommon. Their numerical values need not be identical.
A historical rule combining PRR ≥ 2, chi-square ≥ 4 and at least three reports is widely quoted in pharmacovigilance literature. It should not be presented as a universal current EU regulatory threshold. Organisations may use different algorithms, and Addendum I explicitly emphasises evaluating the performance of the chosen SDA.
Bayesian and Shrinkage Approaches
Simple disproportionality measures can become unstable when case counts are small. Bayesian and empirical-Bayes approaches reduce this instability by shrinking extreme estimates towards a background value when the available information is sparse.
Information Component
The Information Component (IC) is associated with the WHO Programme for International Drug Monitoring and VigiBase. Conceptually, it compares observed and expected reporting on a logarithmic scale while applying Bayesian shrinkage. Its lower credibility bound can be used as part of a screening rule.
The important teaching point is not to treat IC as a different type of causal evidence. Like ROR and PRR, it is a statistical method for identifying unusual reporting patterns. Bayesian shrinkage changes the behaviour of the estimator in sparse data; it does not convert spontaneous reports into incidence data.
Empirical Bayesian methods
Empirical Bayesian approaches such as the Multi-Item Gamma Poisson Shrinker (MGPS) produce shrunk estimates such as the empirical Bayes geometric mean (EBGM). These approaches can reduce instability caused by very small counts and are useful in large-scale screening systems.
Their interpretation depends on the exact model, prior structure and threshold rules. A company using such methods should therefore document the implementation sufficiently for reviewers to understand what is being calculated and how flags are generated.
Worked Example and Interpretation
Consider the following illustrative counts from a spontaneous-report database:
- a = 20 reports containing the product and event of interest;
- b = 980 reports containing the product and other events;
- c = 200 reports containing the event with other products; and
- d = 79,800 reports containing other products and other events.
The ROR is:
ROR = (20 × 79,800) / (980 × 200) ≈ 8.14.
The PRR is:
PRR = [20 / 1,000] ÷ [200 / 80,000] = 8.0.
Both measures show strong disproportionate reporting in this hypothetical dataset. That conclusion should still be phrased carefully. It does not mean that the medicinal product increases the clinical incidence of the event eightfold. The database contains reports, not a defined cohort of treated and untreated patients. Exposure denominators, reporting propensity and confounding are not controlled in the way required for an incidence or relative-risk estimate.
The next step is therefore clinical and scientific review rather than causal declaration.
What should be reviewed after an SDR?
A reviewer may examine:
- the underlying cases and their diagnostic quality;
- temporal relationship between exposure and event;
- dechallenge or rechallenge information where relevant;
- indication and disease-related confounding;
- concomitant medicines and competing causes;
- duplicates and follow-up information;
- seriousness, outcome and clinical specificity;
- whether the reaction is already labelled or otherwise known;
- time trends and reporting stimulated by publicity or regulatory action;
- literature, clinical-trial, epidemiological or mechanistic evidence; and
- whether a class effect or interaction is plausible.
The statistical flag is valuable because it tells the reviewer where to look. The decision about whether the information constitutes a signal depends on the totality of evidence.
Data Preparation and Analytical Choices
Disproportionality results can change materially according to how the dataset is constructed. Methodological transparency therefore matters as much as the formula itself.
Product definition
The analysis may be conducted at the level of a brand, active substance, combination, formulation or route. Combining records too broadly can obscure a formulation-specific problem; separating them too narrowly can fragment a true signal. The product definition should match the scientific question.
Event definition
MedDRA Preferred Terms are commonly used, but some safety concepts require a group of terms rather than one PT. Standardised MedDRA Queries or custom groupings can sometimes better represent syndromes or clinical concepts. Term grouping should be defined before interpretation because changing the case definition can alter the apparent disproportionality.
Duplicate reports
Duplicates inflate counts and can create or amplify an SDR. Duplicate detection and management are therefore important prerequisites for statistical screening. The exact process depends on the database and available identifiers, but the organisation should understand how duplicate management affects the dataset used for analysis.
Included report types
Spontaneous-report databases may contain spontaneous cases, literature cases, solicited reports or other report types. Different sources have different reporting mechanisms and may distort comparison if mixed without consideration. The analytical dataset should therefore be defined and documented.
Comparator background
The comparator is usually the rest of the database, but this is not a neutral control group. Its composition changes with geography, product mix, time and reporting practice. Restricting the comparator can sometimes reduce confounding but can also reduce statistical power or introduce new bias. Any restriction should be scientifically justified.
Bias, Confounding and Masking
Disproportionality analysis is particularly vulnerable to biases that are intrinsic to spontaneous reporting.
Stimulated reporting
Media attention, a regulatory communication, a product launch or a newly recognised adverse reaction can increase the probability that clinicians and patients report a particular event. The resulting increase in ROR or PRR may reflect heightened reporting rather than a change in underlying incidence.
Confounding by indication
A medicine may appear associated with an event because the underlying disease itself causes that event. For example, an oncology treatment may be disproportionately reported with events that are common consequences of advanced cancer. Clinical review and appropriate comparator strategies are needed to interpret such patterns.
Channeling and population differences
Medicines are not prescribed randomly in routine care. One product may be preferentially used in older, sicker or treatment-refractory patients. Differences in the treated population can therefore influence reporting patterns.
Masking and competition
Strongly reported combinations can alter the background against which other combinations are compared. A dominant product-event pair may make another true association appear less disproportionate. This phenomenon is often described as masking or competition bias.
Indication and co-medication effects
The event may be more strongly related to a co-prescribed medicine, treatment combination or clinical indication than to the product under review. Signal assessment should therefore avoid interpreting the product-event table in isolation from the clinical context.
Thresholds and Monitoring Frequency
A threshold converts a continuous statistic into an operational flag. The threshold is useful for prioritisation, but no single value has universal validity.
GVP Module IX Addendum I explains that threshold selection influences both false-positive and false-negative performance. It also notes that monitoring frequency may vary according to the product and accumulation of knowledge. A one-month interval has been studied in validation work, while more frequent monitoring has been used in certain settings, but this should not be turned into an invented universal monthly requirement for every MAH product.
The organisation should be able to explain:
- what statistic and confidence or credibility bound are used;
- whether a minimum count is required;
- whether thresholds differ by product or lifecycle stage;
- how subgrouping is applied;
- what frequency of screening is used; and
- how the method's limitations are addressed in clinical review.
The regulatory objective is a functioning signal-detection system, not conformity with one historical numerical rule.
Governance and Practical Implementation
Disproportionality analysis should be embedded within the wider signal-management process rather than operated as a separate statistical service. The outputs only become useful when responsibility for review, clinical interpretation, documentation and escalation is clear.
A practical operating model usually includes defined data sources, a documented analytical method, reproducible data extraction, controlled code or software configuration, trained reviewers and a signal-tracking process that records what happened to each material flag.
The exact division of work will vary. A statistician or safety scientist may generate the analysis, while a medically qualified reviewer interprets the associated cases. A signal-management forum may prioritise validated signals. The QPPV should have appropriate visibility of significant signal-management matters through the pharmacovigilance governance system but need not personally calculate or approve every ROR.
Reproducibility
A reviewer should be able to reconstruct the result sufficiently to understand the decision. Relevant records may include the data cut-off, product and event definitions, inclusion criteria, software or script version, analysis parameters, output and reviewer conclusion.
Reproducibility does not necessarily require preserving a complete copy of every source database for every run. The evidence model should be proportionate and should preserve what is needed to explain how a material result was generated and interpreted.
Change control
Changes to the algorithm, coding hierarchy, database scope or software can alter SDR generation. Significant changes should therefore be assessed for impact. Depending on the change, retrospective comparison or back-testing may be useful to understand whether previously reviewed combinations would behave differently.
This is recommended risk-based practice rather than a requirement to rerun every historical analysis after every software update.
Potential Failure Modes
The following are illustrative failure modes rather than cited inspection findings.
A statistical flag is labelled a safety signal automatically
An SDR is an analytical output. Treating every SDR as a validated signal creates unnecessary workload and confuses screening with scientific judgement.
A fixed threshold is copied without understanding its performance
Historical PRR or ROR cut-offs can be convenient, but they should not be treated as universally optimal. The false-positive and false-negative behaviour of the signal detection algorithm depends on the database, threshold combination, minimum counts and product portfolio.
Report counts are interpreted as incidence
Spontaneous-report databases normally lack a reliable denominator of exposed patients. The number of reports per product should therefore not be presented as an incidence rate unless an appropriate denominator and study design support that calculation.
The underlying cases are never reviewed
A high ROR driven by duplicates, confounding or low-quality cases may disappear on clinical review. Statistical ranking cannot substitute for examination of the evidence that generated the statistic.
Changes in reporting environment are ignored
Publicity, new indications, regulatory action, launch effects or changes in database composition can move the statistic independently of a biological change in risk.
The analysis cannot be reproduced
If the organisation cannot identify the dataset, coding version, parameters or analytical method used for a material decision, later reviewers may be unable to understand why the decision was made.
Inspection Considerations
An inspector assessing statistical signal detection is likely to focus on whether the method is scientifically appropriate, documented, reproducible and integrated with signal management. Illustrative questions include:
- Which spontaneous-report datasets are screened and why?
- Which disproportionality statistic and SDA are used?
- How were thresholds and minimum counts selected?
- How are EudraVigilance responsibilities integrated into the process?
- How are duplicates, coding and product definitions handled?
- What happens when an SDR is generated?
- Who performs the medical review?
- How are false positives, known reactions and confounding documented?
- Can a material historical output be reconstructed?
- How are algorithm or software changes assessed?
- How are results integrated with literature, clinical and epidemiological evidence?
These are potential inspection questions, not a fixed regulatory checklist.
Practical Checklist
For a mature disproportionality process, the organisation should be able to confirm that:
- the purpose and scope of statistical screening are defined;
- the analytical dataset and included report types are understood;
- product and event definitions are reproducible;
- duplicate management occurs before or as part of analysis where relevant;
- the statistic and SDA are documented;
- thresholds are justified rather than copied uncritically;
- confidence or credibility bounds are interpreted appropriately;
- data cut-offs and analysis versions are traceable;
- SDRs undergo clinical review before regulatory conclusions are drawn;
- spontaneous-report statistics are not described as incidence or causal risk estimates;
- literature and other evidence sources are considered where needed;
- algorithm and software changes are controlled proportionately; and
- material decisions are retained in the signal-management record.
Key Takeaways
Disproportionality analysis is a screening method for spontaneous-report data. It identifies product-event combinations reported more frequently than a database-derived background. ROR, PRR, IC and empirical-Bayes methods express that concept differently, but none of them proves causality or directly measures incidence.
The practical unit of interest is the signal of disproportionate reporting: a combination that meets the organisation's defined signal-detection algorithm and therefore warrants review. The choice of algorithm, thresholds, minimum counts, subgrouping and monitoring frequency affects performance and should be justified for the system in which it is used.
EudraVigilance uses the reporting odds ratio within EVDAS, and EMA's GVP Module IX Addendum I provides the principal EU methodological guidance. Its central message is proportionate statistical screening followed by clinical and scientific evaluation—not automatic action from a numerical threshold.
A robust process therefore links method → reproducible output → clinical review → signal decision. The value of the statistic lies in focusing expert attention where it is most needed.
References
- European Medicines Agency. Guideline on good pharmacovigilance practices (GVP) Module IX – Signal management (Rev. 1). EMA/827661/2011 Rev. 1. https://www.ema.europa.eu/en/documents/scientific-guideline/guideline-good-pharmacovigilance-practices-gvp-module-ix-signal-management-rev-1_en.pdf
- European Medicines Agency. GVP Module IX Addendum I – Methodological aspects of signal detection from spontaneous reports of suspected adverse reactions. EMA/209012/2015. https://www.ema.europa.eu/en/documents/scientific-guideline/guideline-good-pharmacovigilance-practices-gvp-module-ix-addendum-i-methodological-aspects-signal_en.pdf
- European Medicines Agency. EudraVigilance system overview — EVDAS and reporting odds ratio. https://www.ema.europa.eu/en/human-regulatory-overview/research-development/pharmacovigilance-research-development/eudravigilance/eudravigilance-system-overview
- European Medicines Agency. Signal management. https://www.ema.europa.eu/en/human-regulatory-overview/post-authorisation/pharmacovigilance-post-authorisation/signal-management
- European Commission. Commission Implementing Regulation (EU) No 520/2012, consolidated text. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02012R0520-20260212
- Council for International Organizations of Medical Sciences. Practical Aspects of Signal Detection in Pharmacovigilance: Report of CIOMS Working Group VIII. CIOMS, 2010.
Regulatory Note
EU legislation establishes binding pharmacovigilance obligations within its scope. GVP Module IX and Module IX Addendum I provide regulatory guidance for signal management and statistical signal detection. Specific ROR, PRR, IC or Bayesian thresholds, minimum case counts, subgrouping strategies, monitoring frequencies and software implementations are methodological choices unless an applicable regulatory source specifies otherwise. Historical rules such as PRR ≥ 2 with chi-square and minimum-count criteria should not be presented as universal current EU requirements. A signal of disproportionate reporting is a screening output and should not be interpreted as proof of causality, incidence or relative risk.