GVP Module VI: Duplicate Management of Suspected Adverse Reaction Reports
- GVP Module VI: Duplicate Management of Suspected Adverse Reaction Reports
- Introduction
- 1. What Is a Duplicate?
- 2. Why Duplicate Management Matters
- 3. Where Duplicates Come From
- 4. Duplicate Detection Is a Process
- 5. Information Used to Identify Duplicates
- 6. No Single Field Is Sufficient
- 7. Duplicate Detection During Intake
- 8. Duplicate Detection During Follow-Up
- 9. Potential Duplicate Versus Confirmed Duplicate
- 10. Clinical Context Matters
- 11. Reporter Information
- 12. Patient Information
- 13. Systematic Duplicate Searches
- 14. False Positives and False Negatives
- 15. Matching Criteria Should Be Risk-Based
- 16. Clinical Narrative as a Matching Tool
- 17. Reporter Relationships
- 18. Temporal Relationships
- 19. Product Information
- 20. Duplicate Assessment in a Multinational System
- 21. Duplicate Assessment and EudraVigilance
- 22. Confirming a Duplicate
- 23. Preserving Information From Both Reports
- 24. Selecting the Master Case
- 25. Regulatory Reporting After Duplicate Identification
- 26. Follow-Up of a Duplicate
- 27. Duplicate Management and Signal Detection
- 28. Duplicate Metrics
- 29. Common Duplicate-Management Failures
- 30. A Practical Duplicate Decision Tree
- 31. Quality Assurance of Duplicate Management
- 32. Practical Example: Two Reports, One Patient
- 33. Practical Example: Similar Reports, Different Patients
- 34. Practical Example: Follow-Up Reveals a Duplicate
- 35. Practical Example: Different Product Names
- 36. Practical Example: Literature and Spontaneous Report
- 37. Duplicate Management and Personal Data
- 38. Duplicate Management and Data Retention
- 39. Inspection Questions on Duplicate Management
- 40. Audit of Duplicate Management
- 41. Duplicate Management and Signal Detection
- 42. What Good Looks Like
- 43. Final Principles
- Key Takeaways
- References
- Regulatory Note
Introduction
Duplicate management is a fundamental data-quality control in pharmacovigilance.
The same underlying patient experience may be reported through several channels. A healthcare professional may report an event directly to a marketing authorisation holder, the patient may contact a company, a competent authority may transmit information, and a business partner may provide another version of the same report.
If these reports are treated as independent cases, the safety database can overstate the number of patients and events. If they are incorrectly merged, clinically important information can be lost.
Duplicate management therefore requires a balance between two objectives:
Identify reports describing the same case
↓
Prevent double counting
↓
Preserve all meaningful information
↓
Maintain traceability to original sources
GVP Module VI Addendum I provides specific guidance on duplicate management. This article focuses on the practical principles needed to apply that guidance within an operational pharmacovigilance system.
1. What Is a Duplicate?
A duplicate is a report that concerns the same individual case as another report already present in the pharmacovigilance system.
The key concept is same underlying case, not simply similar clinical content.
Two patients taking the same medicinal product and experiencing the same adverse reaction are not duplicates merely because their reports look similar.
Conversely, reports containing substantially different descriptions may still represent the same patient and clinical event.
2. Why Duplicate Management Matters
Duplicate cases can affect multiple parts of pharmacovigilance.
They can distort:
- case counts;
- reporting frequencies;
- signal detection;
- disproportionality analyses;
- aggregate reports;
- risk assessments;
- and regulatory reporting.
The impact is therefore much greater than database housekeeping.
Duplicate management is part of ensuring that safety information represents patients and clinical events accurately.
3. Where Duplicates Come From
Potential duplicates can arise from many sources.
Examples include:
- multiple healthcare professionals reporting the same patient;
- patient and healthcare-professional reports;
- affiliate and central reports;
- partner and MAH reports;
- MAH and competent-authority reports;
- literature and spontaneous reports;
- follow-up information submitted through another channel;
- and transfers between databases.
The greater the number of intake channels, the more important systematic duplicate detection becomes.
4. Duplicate Detection Is a Process
Duplicate detection should not be reduced to a single database search.
A mature process includes:
Potential case
↓
Initial duplicate search
↓
Case processing
↓
Ongoing duplicate assessment
↓
New information received
↓
Reassessment
↓
Confirmed duplicate?
↓
Controlled linking / merging
A case that does not appear to be a duplicate at initial receipt may become recognisable as a duplicate when additional information is received.
5. Information Used to Identify Duplicates
Potential duplicates can be identified using combinations of information such as:
- patient age or age group;
- sex;
- initials or other identifiers;
- country;
- dates of exposure;
- event dates;
- medicinal product;
- dose and route;
- adverse reactions;
- laboratory findings;
- reporter details;
- medical history;
- and distinctive clinical circumstances.
The more distinctive the combination of information, the stronger the basis for assessing whether two reports represent the same case.
6. No Single Field Is Sufficient
Duplicate detection should not depend on one identifier alone.
A patient's age may be identical in two unrelated cases. A common adverse reaction may occur in many patients. A product and date combination may also occur repeatedly.
The assessment should therefore consider the overall evidence.
For example:
Same product
+ same country
+ similar event date
+ same age
+ same unusual laboratory finding
+ same reporter
↓
High suspicion of duplication
The conclusion should still be based on the available evidence and documented according to the organisation's procedure.
7. Duplicate Detection During Intake
The first duplicate check should normally occur when a new report is being processed.
The processor should search for existing cases using the information available at that point.
The search strategy should be sufficiently broad to identify plausible matches without creating an unmanageable number of false positives.
System-supported searches can improve consistency, but automated matching should not automatically determine the final outcome where clinical judgement is required.
8. Duplicate Detection During Follow-Up
Follow-up can reveal information that was not available when the original case was entered.
For example, a follow-up may provide:
- the patient's exact age;
- hospital details;
- reporter identity;
- specific laboratory results;
- treatment dates;
- or a distinctive clinical sequence.
This information can make a previously unrecognised duplicate apparent.
Duplicate assessment should therefore remain part of ongoing case management.
9. Potential Duplicate Versus Confirmed Duplicate
A possible match is not automatically a confirmed duplicate.
The organisation should distinguish between:
- potential duplicate — sufficient similarity exists to require assessment;
- confirmed duplicate — evidence supports that the reports describe the same underlying case.
This distinction is important because prematurely merging uncertain cases can result in loss of legitimate patient information.
10. Clinical Context Matters
Clinical context can be decisive.
Two reports may contain similar product and reaction terms but clearly describe different patients or different episodes.
Conversely, a report with different terminology may describe the same patient if the clinical sequence, reporter and exposure history align.
Duplicate assessment should therefore consider the clinical narrative rather than relying exclusively on coded fields.
11. Reporter Information
Reporter information can be particularly useful in duplicate assessment.
Two reports from the same healthcare professional describing the same unusual event may have a high probability of representing the same case.
However, the same reporter can also report different patients with the same medicinal product and reaction.
Reporter identity is therefore evidence, not proof.
12. Patient Information
Patient characteristics can provide important matching information.
Relevant characteristics can include:
- age;
- sex;
- initials or other permitted identifiers;
- weight;
- relevant medical history;
- hospital or treatment setting;
- and dates associated with exposure or clinical events.
Personal data should be handled according to applicable privacy and pharmacovigilance requirements.
The next chunk will cover duplicate algorithms, clinical assessment, linking and merging, information preservation, follow-up, EudraVigilance and practical examples.
13. Systematic Duplicate Searches
A pharmacovigilance database should support a systematic approach to duplicate detection.
Potential matching can use combinations of structured and unstructured information. Automated algorithms may flag candidate pairs, while trained reviewers determine whether the reports actually represent the same underlying case.
A useful conceptual model is:
Candidate generation
↓
Similarity assessment
↓
Clinical review
↓
Confirmed duplicate?
↙ ↘
Yes No
↓ ↓
Controlled Retain as
link/merge separate case
The algorithm should be regarded as a detection aid, not as a substitute for appropriate assessment.
14. False Positives and False Negatives
Any duplicate-detection system can produce two important types of error.
A false positive occurs when two different cases are incorrectly treated as duplicates.
A false negative occurs when two reports representing the same case are treated as separate cases.
The consequences are different.
False positives can result in loss or distortion of patient information. False negatives can result in double counting and distorted safety analyses.
The process should therefore be designed to minimise both risks.
15. Matching Criteria Should Be Risk-Based
Not every potential match requires the same level of investigation.
A highly distinctive clinical pattern may justify rapid escalation for review, while two reports sharing only a common product and common reaction may require little further action.
The organisation should define appropriate criteria and escalation rules so that duplicate review is consistent across processors and teams.
16. Clinical Narrative as a Matching Tool
Structured fields can miss important similarities.
The narrative may reveal that two reports describe:
- the same admission;
- the same diagnostic investigation;
- the same treatment sequence;
- the same unusual event;
- or the same healthcare professional and clinical setting.
For difficult cases, narrative review can therefore be decisive.
17. Reporter Relationships
Reporter relationships can provide useful supporting evidence.
For example, a physician may report a serious reaction and subsequently a hospital pharmacist may report additional information about the same patient.
The reports may initially appear different because the reporters provide different portions of the clinical history.
The organisation should avoid interpreting differences in wording as evidence that the reports necessarily concern different patients.
18. Temporal Relationships
Dates can be highly informative.
Relevant dates may include:
- treatment initiation;
- treatment discontinuation;
- event onset;
- hospital admission;
- investigation;
- treatment of the reaction;
- and report receipt.
A compatible sequence can strengthen a duplicate assessment, while clearly incompatible timelines may support the conclusion that two cases are distinct.
19. Product Information
Product information should be compared carefully when assessing potential duplicates.
Relevant details can include:
- active substance;
- product name;
- strength;
- formulation;
- route;
- batch information where relevant;
- manufacturer;
- and treatment dates.
Different product names do not necessarily mean different products, particularly where the same active substance is marketed under multiple names.
20. Duplicate Assessment in a Multinational System
Multinational organisations face additional complexity because the same case can appear in several national workflows.
For example:
Patient in Member State A
↓
Local affiliate report
↓
Global safety database
↑
Competent-authority report
↑
Partner report
Each route can create a separate database record before the relationship becomes apparent.
Global duplicate controls should therefore work across affiliates, products, reporters and reporting sources.
21. Duplicate Assessment and EudraVigilance
EudraVigilance can contain reports originating from different organisations and regulatory pathways.
MAHs therefore need processes that take account of duplicate information received through regulatory and other channels.
The organisation should not assume that a report is unique simply because it arrived through a different channel or has a different case identifier.
Case identifiers are important for traceability but are not, by themselves, proof that two reports describe different patients.
22. Confirming a Duplicate
When the available evidence supports that two records represent the same underlying case, the organisation should follow its controlled duplicate-management procedure.
The procedure should define:
- who can make the determination;
- what evidence is required;
- how the relationship is documented;
- how the records are linked or merged;
- and how subsequent follow-up is managed.
The determination should remain auditable.
23. Preserving Information From Both Reports
Confirming a duplicate should not mean deleting useful information.
One report may contain detailed medical history while another contains a laboratory result or outcome information.
The resulting case representation should preserve relevant information from the available sources while maintaining traceability to the original reports.
This is one reason duplicate management requires controlled procedures rather than simple record deletion.
24. Selecting the Master Case
Where a system uses a master or primary case concept, the selection should follow a defined approach.
The choice may depend on factors such as:
- completeness;
- earliest valid receipt;
- source quality;
- regulatory status;
- and system architecture.
The organisation should be able to explain the rationale and ensure that information from other reports is not lost.
25. Regulatory Reporting After Duplicate Identification
Duplicate identification can affect regulatory reporting.
If two reports have already been submitted, the organisation may need to manage the relationship and any subsequent reporting according to the applicable rules and technical requirements.
The correct action depends on the circumstances and the applicable regulatory framework.
Duplicate management should therefore be integrated with regulatory-reporting procedures rather than handled only as an internal database exercise.
26. Follow-Up of a Duplicate
A duplicate can contain information that is absent from the existing case.
The organisation should assess whether the additional information should be incorporated into the retained case and whether it changes:
- seriousness;
- clinical interpretation;
- outcome;
- causality;
- reportability;
- or other relevant assessments.
The fact that a report is a duplicate does not make all information within it irrelevant.
27. Duplicate Management and Signal Detection
Duplicate control is particularly important for signal detection.
If one clinical event appears as multiple independent cases, apparent reporting frequency may increase artificially.
Conversely, incorrect merging of distinct cases can reduce the apparent frequency of an important event.
Duplicate management therefore contributes directly to the reliability of downstream safety analyses.
28. Duplicate Metrics
Useful duplicate-management metrics may include:
- potential duplicates identified;
- confirmed duplicate rate;
- false-positive rate where measurable;
- duplicate detection by source;
- time to duplicate resolution;
- and recurring duplicate patterns.
Trends can reveal weaknesses in intake channels or system interfaces.
For example, a sudden increase in duplicates from one affiliate may indicate a change in local workflow rather than a change in the underlying safety profile.
29. Common Duplicate-Management Failures
Relying on exact matches
Two reports are considered different because one or more fields are not identical.
Relying on one identifier
A shared patient characteristic is treated as sufficient proof of duplication.
Merging too aggressively
Cases are combined because they look similar, resulting in loss of legitimate patient information.
Checking only at initial entry
Later information is not reassessed for potential duplication.
Ignoring narratives
Only structured fields are compared, despite important clinical information being contained in free text.
Treating case IDs as patient IDs
Different case numbers are incorrectly interpreted as evidence of different underlying cases.
30. A Practical Duplicate Decision Tree
A simple operational framework is:
Is there an existing case with plausible overlap?
↓
Yes
↓
Compare patient + product + event + dates + reporter + narrative
↓
Is evidence sufficient?
↙ ↘
Yes No
↓ ↓
Confirmed duplicate Potential duplicate
↓ ↓
Controlled link/merge Further assessment
↓ ↓
Preserve information Reassess when information changes
The exact workflow should reflect the applicable GVP guidance and the organisation's validated systems and procedures.
31. Quality Assurance of Duplicate Management
Duplicate management itself should be subject to quality oversight.
Possible controls include:
- targeted QC samples;
- periodic review of confirmed duplicates;
- assessment of missed duplicates;
- trend analysis;
- system-performance testing;
- and audit of the overall process.
The goal is to determine whether the duplicate process is effective, not simply whether processors followed a checklist.
The next chunk will cover practical duplicate cases, inspection considerations, privacy and masking interfaces, and the final References + Regulatory Note.
32. Practical Example: Two Reports, One Patient
A healthcare professional reports a patient who developed a serious reaction shortly after treatment. A week later, the patient's pharmacist submits a report describing the same reaction, hospitalisation and treatment.
The reports have different case identifiers and slightly different terminology.
The duplicate assessment should compare the patient characteristics, product, dates, clinical course, healthcare setting and reporter information.
If the evidence supports that both reports describe the same clinical episode, they should be managed as one underlying case while preserving the information contributed by both sources.
33. Practical Example: Similar Reports, Different Patients
Two patients receive the same product and both develop the same common adverse reaction within a similar period.
Their ages are similar and both reports originate from the same hospital.
These similarities do not establish duplication.
The organisation should assess the broader evidence, including patient identifiers, dates, clinical details and reporter information.
Where the evidence supports two separate patients, the cases should remain separate.
34. Practical Example: Follow-Up Reveals a Duplicate
A case is initially processed from a consumer report containing limited information.
Several weeks later, follow-up information identifies the treating physician and hospital admission date. A search now identifies another case from the same physician concerning a patient with the same unusual clinical presentation.
This demonstrates why duplicate assessment must continue after initial case creation.
35. Practical Example: Different Product Names
A patient report refers to a branded product while a second report refers to the active substance.
A superficial database search might fail to identify the potential match.
The duplicate process should therefore account for product synonyms, trade names, active substances and other relevant product identifiers.
Product nomenclature controls are an important part of duplicate detection.
36. Practical Example: Literature and Spontaneous Report
A published case describes a patient exposed to a medicinal product and experiencing a serious reaction. A spontaneous report concerning the same patient has already been received.
The literature report should be assessed against the existing case rather than automatically entered as an independent patient.
If the reports are duplicates, the relevant clinical information from the publication should be incorporated appropriately and the source relationship retained.
37. Duplicate Management and Personal Data
Duplicate detection may require comparison of patient and reporter information.
These activities must be performed within the applicable data-protection framework and the organisation's controlled pharmacovigilance processes.
The objective is not to collect unnecessary personal information. It is to use the information legitimately available for pharmacovigilance to determine whether reports represent the same underlying case.
The dedicated article on masking and protection of personal data will address this subject in greater detail.
38. Duplicate Management and Data Retention
When records are linked or merged, the organisation should preserve sufficient history to demonstrate what happened to the original reports.
A reviewer should be able to understand:
- which reports were considered duplicates;
- why the determination was made;
- which case was retained as the principal record where applicable;
- what information was incorporated;
- and how subsequent updates were managed.
This historical traceability is important for both data integrity and inspection readiness.
39. Inspection Questions on Duplicate Management
An inspector may ask:
- How do you identify potential duplicates?
- What matching criteria are used?
- Is duplicate detection automated?
- Who makes the final determination?
- How are uncertain cases handled?
- How do you reassess duplicates after follow-up?
- How are duplicates received from different affiliates or partners managed?
- How do you prevent loss of information when records are merged?
- How do you monitor the effectiveness of duplicate detection?
- Can you demonstrate a recent duplicate assessment?
The strongest evidence is an ordinary case history showing that the process works as designed.
40. Audit of Duplicate Management
Duplicate management can itself be audited.
An audit may examine:
- duplicate procedures;
- system configuration;
- search strategies;
- reviewer training;
- samples of confirmed duplicates;
- samples of cases initially considered distinct;
- missed duplicates;
- metrics;
- and CAPA.
The audit should test whether the control actually prevents material data-quality problems.
41. Duplicate Management and Signal Detection
A mature pharmacovigilance system should understand how duplicate decisions affect downstream analytics.
If duplicates are systematically over-counted, signals may appear stronger than the underlying evidence supports.
If distinct cases are systematically merged, signals may be weakened.
Duplicate management is therefore one component of the broader data-quality chain supporting signal detection and benefit-risk assessment.
42. What Good Looks Like
A mature duplicate-management process has:
- defined criteria;
- systematic searches;
- appropriate use of automated detection;
- trained human assessment;
- clinical review for difficult cases;
- controlled linking or merging;
- preservation of source information;
- reassessment following follow-up;
- multinational and partner interfaces;
- appropriate metrics;
- and quality oversight.
The process should be sufficiently sensitive to identify plausible duplicates without aggressively combining unrelated cases.
43. Final Principles
- A duplicate is a report concerning the same underlying individual case, not simply a similar report.
- Duplicate detection should begin at case intake and continue throughout the case lifecycle.
- Multiple data elements should be assessed together.
- Structured fields and clinical narratives both provide important evidence.
- Automated matching is a useful detection aid but should not replace appropriate assessment.
- Potential duplicates and confirmed duplicates should be distinguished.
- Confirmed duplicates should be managed through controlled procedures.
- Important information from all source reports should be preserved.
- Different case identifiers do not necessarily mean different patients.
- Duplicate management should work across affiliates, partners, literature and regulatory sources.
- Duplicate decisions can materially affect signal detection and aggregate safety analyses.
- The effectiveness of duplicate management should itself be monitored and periodically assessed.
Key Takeaways
- Duplicate management is a substantive pharmacovigilance data-quality control.
- The question is whether reports describe the same underlying patient and clinical episode.
- Similarity alone does not prove duplication, and differences in wording do not prove that reports are distinct.
- Duplicate detection should use multiple data elements and appropriate clinical judgement.
- Follow-up can reveal duplicates that were not recognisable initially.
- Confirmed duplicates must be managed without losing clinically meaningful information or traceability.
- The process should operate across the complete multinational pharmacovigilance network.
- Duplicate management directly affects the reliability of downstream safety analyses.
References
- European Medicines Agency. Good Pharmacovigilance Practices (GVP), Module VI — Collection, management and submission of reports of suspected adverse reactions to medicinal products. Current version should be consulted for the overarching ICSR framework.
- European Medicines Agency. GVP Module VI — Addendum I: Duplicate management of suspected adverse reaction reports. Primary EU guidance for duplicate management.
- European Medicines Agency. GVP Module VI — Addendum II: ICSR masking of personal data. Relevant to the privacy and personal-data interface.
- European Medicines Agency. EudraVigilance guidance and electronic reporting requirements. Relevant to duplicate information received through the EU electronic reporting environment.
- International Council for Harmonisation. ICH E2B(R3) — Individual Case Safety Reports. Relevant to structured electronic ICSR data.
- European Parliament and Council. Directive 2001/83/EC, as amended. EU legal framework for medicinal products for human use and pharmacovigilance.
- European Parliament and Council. Regulation (EC) No 726/2004, as amended. Union framework for authorisation and supervision of medicinal products and relevant pharmacovigilance obligations.
Regulatory Note
This article is an educational and practical explanation of duplicate management under the EU GVP framework. It does not replace the current GVP Module VI guidance, Addendum I, applicable EU legislation, EudraVigilance requirements or organisation-specific procedures.
GVP guidance and technical requirements may change. Before applying this article to a live case or regulatory submission, verify the current EMA guidance, applicable legislation, technical specifications and effective dates.
The practical examples are illustrative and are not descriptions of specific regulatory inspection cases unless an authoritative source is explicitly identified.