Data Sources for Signal Management
A safety signal is rarely understood from a single data stream. One source may first reveal an unusual pattern, another may clarify whether the pattern is biologically plausible, and a third may provide the denominator or comparator needed to estimate whether the observed risk differs from what would otherwise be expected. Signal management is therefore an exercise in evidence integration, not simply database screening.
This is especially important because pharmacovigilance data sources are heterogeneous. Spontaneous reports are sensitive to rare and unexpected events but are affected by under-reporting and reporting bias. Randomised trials provide controlled comparisons but may be too small or too selective to detect uncommon post-authorisation risks. Healthcare databases can support comparative epidemiology but introduce confounding, coding limitations and data-provenance questions. Scientific literature may provide rich clinical detail while also reflecting publication bias. None of these limitations makes a source unusable; they determine what questions that source can answer reliably.
GVP Module IX therefore frames signal management as a structured process that draws on all appropriate sources of information. The professional task is to understand what each source contributes, what it cannot establish on its own, and how evidence from different sources changes the level of confidence in a safety hypothesis.
- Data Sources for Signal Management
- Purpose and Regulatory Context
- A Framework for Understanding Signal Data Sources
- Spontaneous Reports
- Scientific Literature
- Clinical Trial and Development Data
- Epidemiological Evidence
- Registries
- Real-World Data and Real-World Evidence
- Post-Authorisation Safety Studies
- Regulatory Information and Product-Level Context
- Non-Clinical and Mechanistic Evidence
- Integrating Evidence Across Sources
- Governance and Data Quality
- Potential Failure Modes
- Disproportionality is treated as proof
- Report counts are interpreted without exposure context
- Only confirming evidence is collected
- Literature searching is too narrow
- Real-world data are assumed to be self-validating because they are large
- Global sources and local knowledge are disconnected
- The signal file cannot reconstruct the analysis
- Inspection Considerations
- Practical Evidence-Integration Example
- Practical Checklist
- Key Takeaways
- References
- Regulatory Note
Purpose and Regulatory Context
The purpose of using multiple signal data sources is to detect and evaluate new or changed risks as early and accurately as possible while avoiding conclusions that the evidence does not support. In the EU, the legal framework for signal management is established through Directive 2001/83/EC, Regulation (EC) No 726/2004 and Commission Implementing Regulation (EU) No 520/2012. GVP Module IX provides the principal operational guidance for signal management within the EU pharmacovigilance system.
GVP Module IX describes signal management as a set of activities including signal detection, validation, confirmation, analysis and prioritisation, assessment, and recommendation for action. Different information sources contribute at different stages, and the same source may contribute differently depending on the medicinal product and safety question.
The role of EudraVigilance
For medicinal products authorised in the EU, EudraVigilance is a central source of suspected adverse reaction data. GVP Module IX and its Addendum I describe the use of EudraVigilance data for signal detection, including statistical screening of spontaneous reports. Marketing authorisation holders have defined responsibilities for monitoring EudraVigilance data for substances within the applicable EU arrangements.
Statistical disproportionality is a detection aid, not a diagnosis of causality. A disproportionality measure asks whether a particular drug-event combination is reported more frequently than a comparison pattern would suggest. It does not establish incidence, eliminate reporting bias or prove that the medicinal product caused the event.
Signal detection is broader than disproportionality
A signal may arise from a small number of clinically persuasive cases, a trial imbalance, an epidemiological study, literature findings, a regulator's assessment, a class effect, a toxicological observation or a combination of these. Conversely, a strong statistical association in a spontaneous-report database may fail validation when case quality, confounding, duplicate reports or stimulated reporting are examined.
The choice of source should therefore follow the scientific question rather than a fixed hierarchy in which one database is assumed to be definitive.
A Framework for Understanding Signal Data Sources
Each source can be evaluated across five dimensions:
- clinical detail — how much information is available about the patient, exposure, event and alternative explanations;
- population context — whether the source provides a denominator or comparison group;
- timeliness — how rapidly new information becomes visible;
- bias and confounding — which systematic distortions are likely; and
- traceability and reproducibility — whether the analysis can be reconstructed from source to conclusion.
These dimensions explain why sources complement rather than replace one another.
| Source type | Typical contribution | Principal limitation |
|---|---|---|
| Spontaneous reports | early recognition of rare, serious or unusual patterns | no reliable exposure denominator; reporting bias |
| EudraVigilance and other spontaneous-report repositories | broad case aggregation and statistical screening | same intrinsic limitations as spontaneous reports plus database heterogeneity |
| Scientific literature | detailed cases, clinical context, mechanistic and epidemiological evidence | publication and selection bias; variable quality |
| Clinical trials | controlled exposure, defined populations and comparator data | limited size, duration and generalisability |
| Observational epidemiology | incidence/rate estimates and comparative risk evaluation | confounding, misclassification and data-quality limitations |
| Registries | longitudinal follow-up in defined populations | selection and completeness issues |
| Real-world healthcare data | large-scale routine-care evidence | coding, missingness, confounding and provenance issues |
| Non-clinical/mechanistic evidence | biological plausibility and hypothesis refinement | uncertain translation to clinical risk |
| Regulatory and product information | external assessments and known-risk context | may be retrospective or jurisdiction-specific |
The table is a scientific orientation, not a mandatory regulatory ranking.
Spontaneous Reports
Spontaneous reports remain one of the most important sources for identifying previously unrecognised adverse reactions, particularly when the event is rare, clinically distinctive, temporally associated with exposure or occurs in a pattern that is difficult to explain by background disease.
Their strength lies in reach. Once a product is widely used, reports can arise from patient groups, healthcare settings and combinations of comorbidity that were poorly represented in clinical development.
Their weakness is that reporting is not systematic. The number of reports is influenced by awareness, publicity, regulatory action, time since launch, country, indication, seriousness and many other factors. The absence of reports does not prove absence of risk, and report counts alone generally cannot be used to calculate incidence.
Case-level clinical review
The value of a spontaneous report depends heavily on case quality. Signal assessment may consider temporal relationship, dechallenge and rechallenge, dose relationship, relevant investigations, alternative causes, concomitant medicines, biological plausibility and consistency across cases.
No single causality criterion is decisive. A coherent case series can increase confidence, but the purpose of validation is not to force each case into a binary causal classification. It is to decide whether the information collectively supports a new potentially causal association or a new aspect of a known association that merits further evaluation.
Duplicates and follow-up
Duplicate reports can distort case counts and disproportionality analyses. GVP Module VI and its duplicate-management addendum address duplicate handling in ICSR systems. Signal processes should therefore use appropriately deduplicated data and retain traceability to the underlying cases.
Follow-up can materially change the interpretation of a case by clarifying diagnosis, timing, exposure or alternative explanations. The amount of follow-up sought should be proportionate to the clinical and regulatory importance of the missing information rather than driven by an invented universal checklist.
Scientific Literature
Scientific literature can contribute to signal management in several ways. A case report may contain clinical detail unavailable from a spontaneous report. A case series may suggest a recurring phenotype. A pharmacological paper may provide a mechanism that makes an association more plausible. An observational study may quantify risk in a defined population, while a systematic review can place individual findings into a wider evidence base.
The evidentiary value depends on study design and quality. A published case report can be compelling when the event is highly specific and the temporal pattern is persuasive, but it cannot provide a population rate. A large observational study may provide comparative risk estimates but still be vulnerable to confounding or outcome misclassification.
Literature surveillance for ICSR identification and literature use in signal evaluation are related but not identical. Routine literature monitoring supports case collection obligations, whereas signal assessment may require a broader and more question-specific search strategy. The latter should be sufficiently documented that another reviewer can understand what was searched, what evidence was included and how contradictory findings were handled.
Clinical Trial and Development Data
Clinical trials contribute structured exposure and systematically collected safety data. Randomisation, where present, can reduce confounding and make comparisons between treatment groups more interpretable. Trial datasets can therefore be particularly useful when a suspected adverse event is common in the underlying disease and spontaneous reports provide little denominator context.
At the same time, trials may exclude older patients, people with important comorbidities, pregnant individuals or those receiving interacting medicines. Sample sizes and follow-up periods may also be insufficient for rare or delayed adverse reactions. A lack of imbalance in trials therefore does not automatically refute a post-authorisation signal.
Signal assessment may examine pooled trial data, event timing, dose-response patterns, laboratory changes, adjudicated outcomes, discontinuations and subgroup findings. The analysis should respect the original study design and avoid creating apparently precise conclusions from small post hoc subsets.
Development data can also provide an important historical baseline. A post-authorisation event may look new until early-phase pharmacology, non-clinical findings or previously isolated trial cases are revisited in light of the new hypothesis.
Epidemiological Evidence
Epidemiology becomes especially important when the safety question moves from could this association exist? to how large is the risk, in whom, and compared with what? Cohort, case-control, self-controlled and other observational designs can provide quantitative evidence using healthcare databases, registries or purpose-built data collection.
The design should be chosen for the question. A comparative cohort may be useful for estimating relative and absolute rates, while a self-controlled design can reduce confounding by characteristics that do not change over time. No design is automatically superior; each introduces its own assumptions and vulnerabilities.
Confounding and bias
Confounding by indication is a recurring problem in medicine-safety studies because the disease for which a medicine is prescribed may itself increase the risk of the outcome. Channeling, differential healthcare contact, incomplete capture of over-the-counter exposure, diagnostic surveillance and coding practices can also alter results.
A credible assessment therefore describes the data source, exposure definition, outcome definition, comparator choice, covariates, handling of missing data, follow-up, sensitivity analyses and residual uncertainty. Reproducibility matters as much as statistical significance.
Incidence and background rates
Reliable denominator-based data can help distinguish a true increase in risk from an event that is expected to occur frequently in the treated population. Background rates can be useful context, but they should be matched as closely as possible to age, sex, disease, geography, calendar period and outcome definition. Comparing spontaneous-report counts with an unrelated background rate can create a misleading impression of precision.
Registries
Registries can provide longitudinal information in populations that are otherwise difficult to study, including people with rare diseases, pregnancy exposures or patients receiving specialised therapies. Their strength lies in structured follow-up and clinically rich data.
Their limitations include incomplete enrolment, loss to follow-up, variable representativeness and changing data quality over time. A registry designed for disease natural history may also not collect the exposure or outcome variables needed for a particular safety question.
When registry data support signal assessment, the reviewer should understand how patients enter the registry, how outcomes are captured and validated, how missing data are handled and whether the registry population differs materially from the target population for the medicine.
Real-World Data and Real-World Evidence
Real-world data can include electronic health records, claims, dispensing records, laboratory systems and linked healthcare datasets. Real-world evidence is the clinical evidence generated from analysis of such data.
These sources can expand signal assessment beyond spontaneous reporting by providing large denominators, longitudinal follow-up and comparator groups. They are particularly useful for common outcomes, risk factors, patterns of use and comparative analyses that cannot be answered from case reports alone.
However, volume does not guarantee validity. Large datasets can produce precise estimates of a biased quantity if exposure, outcome or confounding variables are poorly measured.
Data provenance and fitness for use
Before interpreting an RWD analysis, the pharmacovigilance team should understand where the data came from, why they were collected, how frequently they are updated, which patients are represented, what transformations have occurred and which variables are reliable for the intended analysis.
A dataset suitable for identifying prescriptions may be unsuitable for confirming a clinical diagnosis. A claims database may capture hospitalisation accurately while lacking laboratory detail. An electronic health record may contain rich narrative information but incomplete records from care delivered outside the network.
This is the practical meaning of fitness for use: data quality is not an abstract score; it is judged against the specific safety question.
Analytical reproducibility
Signal-related RWD analyses should retain sufficient information to reconstruct cohort definitions, code lists, time windows, statistical methods and data versions. Reproducibility enables peer review, sensitivity analysis and later reassessment when the signal evolves.
Post-Authorisation Safety Studies
Post-authorisation safety studies (PASS) may be imposed, required or conducted voluntarily to characterise a safety concern, quantify a risk, identify risk factors or evaluate risk-minimisation measures. GVP Module VIII provides the EU framework for non-interventional PASS within its scope.
PASS should not be treated as a generic data source that automatically confirms a signal. Their evidentiary value depends on the protocol, design, population, endpoint validity, execution and analysis. A study designed to evaluate a risk-minimisation process may contribute little to estimating causal association, while a well-designed comparative database study may directly address the magnitude of risk.
Signal governance should therefore use the study for the question it was designed to answer.
Regulatory Information and Product-Level Context
Regulatory sources can alter the interpretation of a signal even when they do not contain new raw data. These sources include assessment reports, PRAC recommendations, authority requests, safety communications, changes in product information and regulatory action on related products.
A regulatory decision concerning another product in the same pharmacological class may justify reassessing class plausibility, but it should not be imported automatically into the product under review. Differences in molecular structure, target selectivity, formulation, exposure and indication may materially change risk.
The current product information, RMP, previous signal assessments, periodic safety reports and relevant regulatory correspondence provide essential context because they show what is already known, which risks have already been evaluated and whether the apparently new information actually changes the established safety profile.
Non-Clinical and Mechanistic Evidence
Non-clinical toxicology, pharmacology, receptor biology, genetic evidence and mechanistic studies can support or weaken biological plausibility. Their value is often greatest when clinical observations are sparse but coherent with a known pathway.
Mechanistic plausibility should not be treated as proof of clinical causality. Many biological effects observed in vitro or in animals do not translate into clinically meaningful human risk. Conversely, lack of a known mechanism does not invalidate a strong clinical association. Mechanistic evidence is one component of the total evidence, not a gatekeeper.
Integrating Evidence Across Sources
Evidence integration is not a vote in which each source receives one point. Sources answer different questions and vary in reliability for the issue being assessed.
A practical assessment can ask:
- Is the clinical phenotype specific and consistently described?
- Is there a plausible temporal relationship?
- Are alternative explanations sufficient?
- Is there evidence of dose-response, dechallenge or rechallenge?
- Do controlled or comparative data show an imbalance?
- Does epidemiology support an increased rate?
- Is the association observed across settings or populations?
- Is there biological or class plausibility?
- Could reporting, selection, detection or publication bias explain the pattern?
- What evidence would most efficiently reduce the remaining uncertainty?
This sequence helps the team move from source collection to scientific reasoning.
Governance and Data Quality
A signal-management process should make clear which data sources are routinely monitored, which sources are used when a hypothesis requires deeper assessment, who is responsible for obtaining and analysing them, and how the resulting evidence is documented. Governance should support scientific flexibility without allowing undocumented ad hoc analysis.
Source inventory and access
A source inventory can identify the major internal and external data streams available to the organisation, their owners, access arrangements, refresh characteristics and known limitations. This is recommended practice rather than a prescribed GVP template. Its purpose is to prevent important sources from being overlooked and to make dependencies visible.
For external databases and vendors, the organisation should understand contractual access, provenance, transformations, data-quality controls and the conditions under which data may be used. A prestigious vendor name is not a substitute for understanding whether the dataset is fit for the safety question.
Quality controls should match the source
Different sources require different controls. Spontaneous reports require reliable case processing, duplicate management and traceability. Literature searches require reproducible search logic and documented screening. Epidemiological analyses require transparent cohort definitions, code lists and methods. Regulatory intelligence requires reliable identification of relevant decisions and linkage to internal assessment.
A universal checklist applied identically to every source can create paperwork without improving evidence quality.
Documentation of the evidence trail
A mature signal record should allow a reviewer to move from the initial information through validation and assessment to the final recommendation. This does not mean every raw dataset must be copied into the signal file. It means that the source, version or extraction date, analysis method, important outputs, reviewer interpretation and decision rationale should be traceable.
Where analyses are repeated as new data accumulate, versioning should make clear which evidence supported each decision at the time it was made.
Potential Failure Modes
The following are illustrative failure modes rather than cited inspection findings.
Disproportionality is treated as proof
A statistical signal can focus attention efficiently, but it does not establish incidence or causality. The clinical content of the cases and other evidence sources still require review.
Report counts are interpreted without exposure context
Comparing raw numbers between products, countries or time periods can be misleading because reporting intensity and exposure differ. Where denominator information is unavailable, conclusions should reflect that limitation.
Only confirming evidence is collected
Once a plausible hypothesis emerges, reviewers can become vulnerable to confirmation bias. A sound assessment actively looks for contradictory evidence, alternative explanations and analyses that could falsify the working hypothesis.
Literature searching is too narrow
Searching only the exact product-event phrase may miss class effects, synonyms, mechanistic evidence or epidemiological studies using different terminology. Search strategy should expand appropriately as the question evolves.
Real-world data are assumed to be self-validating because they are large
Large sample size reduces random error but does not remove confounding, misclassification or selection bias. Data provenance and fitness for use remain fundamental.
Global sources and local knowledge are disconnected
A global database may not capture an important local regulatory action, utilisation pattern or study. Conversely, an affiliate may identify information with wider relevance but fail to escalate it. Signal governance should connect local and global evidence flows.
The signal file cannot reconstruct the analysis
If query parameters, data cut-off, case series, literature strategy or code lists cannot be retrieved, later reviewers may be unable to understand why a decision was made. Traceability should be designed into the process rather than reconstructed before inspection.
Inspection Considerations
An inspector evaluating signal data sources may examine whether the organisation has a systematic process for identifying, accessing, analysing and integrating relevant safety information. The focus is likely to be on effectiveness and traceability rather than on whether the company uses a specific commercial tool or statistical threshold.
Potential inspection questions include:
- Which sources are routinely monitored for this product and why?
- How are EudraVigilance monitoring responsibilities identified and performed?
- How are duplicate reports controlled before signal analysis?
- How is literature used differently for case collection and for signal assessment?
- What triggers the use of epidemiology or real-world data?
- How is the fitness of an external dataset assessed for a particular question?
- How are analyses versioned and reproduced?
- How are local regulatory information and global signal governance connected?
- How are contradictory findings handled?
- Can the organisation trace a signal decision to the underlying case series, searches, analyses and regulatory context?
- How are important changes in data sources, algorithms or coding versions assessed for impact?
These are illustrative questions; the actual inspection scope is determined by the regulatory authority and inspection context.
Practical Evidence-Integration Example
Consider a hypothetical product for which several spontaneous reports describe acute kidney injury shortly after initiation. A disproportionality analysis also shows an elevated reporting pattern.
The spontaneous reports establish the initial hypothesis, but the next questions require different sources. Clinical review asks whether the cases share timing, dose, dehydration, concomitant nephrotoxic medicines or laboratory patterns. Trial data are reviewed for renal adverse events and creatinine changes. Literature is searched for similar cases, class effects and mechanistic evidence. Exposure data help place the reporting pattern in context. A healthcare-database study may then compare acute kidney injury rates with an appropriate alternative therapy while controlling for renal disease and other confounders.
If the epidemiological estimate does not show increased risk, the signal is not automatically closed. The team asks whether the study captured the specific phenotype, whether exposure and outcome definitions were valid, whether susceptible subgroups were diluted and whether spontaneous reporting was stimulated. Equally, if the observational study shows an association, residual confounding and outcome validation still need consideration.
The final assessment is therefore stronger than any individual result because it explains how the different sources fit together and where uncertainty remains.
Practical Checklist
For a significant signal assessment, the team should be able to confirm that:
- the initial source and data cut-off are identifiable;
- relevant spontaneous reports have been clinically reviewed and appropriately deduplicated;
- statistical screening is interpreted as hypothesis-generating rather than causal proof;
- literature searching is proportionate to the question and reproducible;
- relevant clinical-development data have been considered;
- denominator-based or comparative evidence has been sought where it can materially answer the question;
- epidemiological designs and limitations are understood rather than reduced to effect estimates alone;
- real-world datasets have been assessed for provenance and fitness for use;
- regulatory actions and class information have been evaluated without automatic extrapolation;
- mechanistic evidence is integrated at an appropriate evidentiary level;
- contradictory evidence and alternative explanations are documented;
- important analysis parameters and versions are retrievable;
- the decision record explains what evidence changed the level of concern; and
- remaining uncertainty and proposed next steps are explicit.
Key Takeaways
Signal-management data sources are complementary. Spontaneous reports are often excellent for detecting rare or unusual patterns but generally lack reliable denominators. Clinical trials provide structured comparisons but may be too limited to reveal uncommon or delayed post-authorisation risks. Epidemiological and real-world data can quantify associations in larger populations, but only when confounding, misclassification and data provenance are handled appropriately. Literature, regulatory information and mechanistic evidence add clinical and scientific context.
GVP Module IX does not reduce signal management to a single statistical method. It requires a structured process in which available information is detected, validated, assessed and acted upon. EudraVigilance is a central EU source, and its statistical tools can identify reporting patterns, but disproportionality is one input into clinical and scientific evaluation.
The strongest signal assessments are reproducible evidence arguments. They show what each source contributes, what its limitations are, how contradictory evidence was handled and why the integrated evidence supports a particular conclusion or next step.
The governing question is therefore not which data source is best? but which combination of sources can answer the safety question with the least avoidable uncertainty?
References
- European Medicines Agency. Guideline on good pharmacovigilance practices (GVP) Module IX – Signal management (Rev. 1). EMA/827661/2011 Rev. 1. https://www.ema.europa.eu/en/documents/scientific-guideline/guideline-good-pharmacovigilance-practices-gvp-module-ix-signal-management-rev-1_en.pdf
- European Medicines Agency. GVP Module IX Addendum I – Methodological aspects of signal detection from spontaneous reports of suspected adverse reactions. EMA/209012/2015. Available from the EMA GVP collection: https://www.ema.europa.eu/en/human-regulatory-overview/post-authorisation/pharmacovigilance-post-authorisation/good-pharmacovigilance-practices-gvp
- European Medicines Agency. Questions and answers on signal management, current revision available from EMA signal-management pages. https://www.ema.europa.eu/en/human-regulatory-overview/post-authorisation/pharmacovigilance-post-authorisation/signal-management
- European Medicines Agency. Guideline on good pharmacovigilance practices (GVP) Module VI – Collection, management and submission of reports of suspected adverse reactions to medicinal products (Rev. 2) and Addendum I on duplicate management. Available from the EMA GVP collection.
- European Medicines Agency. Guideline on good pharmacovigilance practices (GVP) Module VIII – Post-authorisation safety studies and applicable addenda. Available from the EMA GVP collection.
- European Commission. Commission Implementing Regulation (EU) No 520/2012, Chapter III provisions on signal management, consolidated text. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02012R0520-20260212
- European Parliament and Council. Directive 2001/83/EC, as amended. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02001L0083
- European Parliament and Council. Regulation (EC) No 726/2004, as amended. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02004R0726
- Council for International Organizations of Medical Sciences. Practical Aspects of Signal Detection in Pharmacovigilance: Report of CIOMS Working Group VIII. CIOMS, 2010.
Regulatory Note
EU legislation establishes binding pharmacovigilance and signal-management obligations within its scope. GVP Module IX and its Addendum I provide the principal EU regulatory guidance used in this article, while GVP Modules VI and VIII govern related ICSR and PASS processes within their respective scopes. Statistical thresholds, literature-review models, epidemiological methods, source inventories, checklists and inspection questions described here are scientific or operational practices unless an authoritative source specifically makes them mandatory. No single data source or disproportionality result should be presented as regulatory proof of causality; signal assessment requires interpretation of the totality of relevant evidence.