Signal Evaluation: DPP-4 Inhibitors and Bullous Pemphigoid
- Signal Evaluation: DPP-4 Inhibitors and Bullous Pemphigoid
- 1. Define the Signal
- 2. Define the Clinical Event Precisely
- 3. Why This Event Was Interesting
- 4. Start With the Background Risk
- 5. The First Signal: Individual Cases
- 6. What Makes an Individual Case More Informative?
- 7. Dechallenge: Useful but Not Definitive
- 8. Rechallenge
- 9. Disproportionality Analysis
- 10. The Vildagliptin Signal
- 11. What Does "Highest Disproportionality" Mean?
- 12. Evidence From Case-Control Studies
- 13. Comparative Cohort Evidence
- 14. Why the Absolute Risk Matters
- 15. Why the Comparator Matters
- 16. New-User Designs
- 17. Propensity Scores
- 18. Age as a Major Confounder
- 19. Neurological Disease
- 20. Clinical Phenotype as Causal Evidence
- 21. Biological Plausibility
- 22. Autoantibodies and BP180
- 23. Why a Class Effect Is Difficult to Establish
- 24. The Importance of Consistency
- 25. A Systematic Review and Meta-Analysis
- 26. Heterogeneity Is Information
- 27. Absolute Risk Versus Relative Risk
- 28. Regulatory Interpretation
- 29. Why "Frequency Not Known" Is Meaningful
- 30. What the Regulatory Action Did Not Establish
- 31. What a Signal Evaluator Should Ask at Each Stage
- 32. A Worked Individual Case
- 33. Why the Individual Case and Population Signal Must Be Connected
- 34. Common Analytical Errors
- Error 1 — Treating disproportionality as incidence
- Error 2 — Treating a case report as proof
- Error 3 — Ignoring background disease risk
- Error 4 — Assuming class effect automatically
- Error 5 — Treating a positive epidemiological result as definitive
- Error 6 — Treating a negative study as definitive reassurance
- Error 7 — Ignoring absolute risk
- Error 8 — Overstating mechanism
- Error 9 — Ignoring phenotype
- Error 10 — Assuming regulatory wording equals causal certainty
- 35. What Is Established, What Is Supported and What Remains Uncertain?
- 36. Regulatory Significance Versus Causal Certainty
- 37. Risk Minimisation
- 38. Why the Signal Matters for Aggregate Safety Review
- 39. A Practical Evidence Table for Signal Review
- 40. Lessons for Signal Management
- Lesson 1 — Start with a precise case definition
- Lesson 2 — Separate signal detection from signal assessment
- Lesson 3 — Use the right denominator
- Lesson 4 — Use comparative designs where possible
- Lesson 5 — Examine absolute as well as relative risk
- Lesson 6 — Investigate confounding explicitly
- Lesson 7 — Examine class versus molecule-specific patterns
- Lesson 8 — Use phenotype as supporting evidence
- Lesson 9 — Explain mechanisms carefully
- Lesson 10 — Do not wait for perfect certainty before considering risk minimisation
- Lesson 11 — Treat regulatory wording as evidence of regulatory significance
- Lesson 12 — Keep reassessing the signal
- 41. What This Case Teaches a QPPV
- Is the event medically credible?
- Is the association biologically plausible?
- Is the signal supported beyond spontaneous reports?
- Is the risk precisely quantified?
- Are there important alternative explanations?
- Is there regulatory significance?
- Does a positive signal mean every case is caused by the medicine?
- Does uncertainty mean the signal should be ignored?
- 42. Overall Signal Assessment
- 43. Final Perspective
- References
Introduction
Some of the most instructive pharmacovigilance signals begin with a small number of apparently unrelated individual cases.
A patient develops an unusual adverse event.
Another patient develops the same event while receiving the same medicine.
Additional cases appear.
The event is uncommon in the general population, the clinical phenotype is recognisable, and the temporal relationship is plausible.
At that point, the pharmacovigilance question is not whether causality has been established.
It has not.
The question is whether the pattern is sufficiently unusual and sufficiently coherent to justify structured investigation.
The historical association between dipeptidyl peptidase-4 (DPP-4) inhibitors and bullous pemphigoid provides a particularly useful example of this process.
Bullous pemphigoid (BP) is an autoimmune subepidermal blistering disease that predominantly affects older adults. It can cause extensive blistering, pruritus, mucosal involvement in some patients, secondary infection and substantial morbidity.
The signal is interesting because it developed through several different evidence streams:
- spontaneous case reports;
- pharmacovigilance database disproportionality;
- clinical case series;
- case-control studies;
- cohort studies;
- comparative pharmacoepidemiology;
- clinical phenotype observations;
- immunological investigations;
- and regulatory assessment.
The evidence did not arrive simultaneously.
Nor did every study produce the same result.
The signal therefore provides a practical demonstration of how a pharmacovigilance team should move from:
case → pattern → signal → epidemiological assessment → regulatory interpretation → risk minimisation.
It also demonstrates an important principle:
A signal can become regulatory-relevant before every aspect of its mechanism or magnitude has been resolved.
1. Define the Signal
The signal can be expressed as:
DPP-4 inhibitor exposure — bullous pemphigoid.
The drug class includes several medicines used in the management of type 2 diabetes mellitus.
Examples include:
- vildagliptin;
- sitagliptin;
- saxagliptin;
- linagliptin;
- and alogliptin.
The signal initially became particularly prominent with vildagliptin.
This distinction matters.
A pharmacovigilance evaluator should not automatically convert:
Drug A → event
into:
Drug class → event
without examining whether the evidence is consistent across individual active substances.
Conversely, an apparent concentration of cases with one molecule does not automatically establish that the risk is restricted to that molecule.
The correct question is:
What does the evidence support at class level, and what does it support for individual active substances?
2. Define the Clinical Event Precisely
Bullous pemphigoid is not simply "rash" or "blistering."
It is an autoimmune blistering disorder characterised by subepidermal separation and autoantibodies directed against components of the basement membrane zone, particularly BP180 and BP230.
Clinical manifestations commonly include:
- tense blisters;
- pruritus;
- urticarial or erythematous lesions;
- erosions following blister rupture;
- and, in some patients, mucosal involvement.
Diagnosis may require:
- clinical assessment;
- skin biopsy;
- direct immunofluorescence;
- serological testing;
- and specialist dermatological evaluation.
This distinction is important in signal evaluation.
If the case definition is too broad, nonspecific blistering disorders may be mixed with true BP.
If it is too narrow, genuine cases may be missed.
Therefore, the evaluator should establish what constitutes a sufficiently specific phenotype before interpreting case counts.
3. Why This Event Was Interesting
Several features made the signal worth investigating.
First, BP is uncommon.
A rare event occurring repeatedly in association with a particular medicine can be more informative than a common nonspecific symptom, provided the diagnosis is reasonably specific.
Second, the patients were often older adults.
Age itself is an important consideration because BP is strongly associated with older age.
Third, diabetes is common in the population receiving DPP-4 inhibitors.
Therefore, the underlying disease and age distribution create an important background-risk problem.
Fourth, BP is an immune-mediated disease.
That raised a biologically interesting question:
Could modification of DPP-4 activity influence immune pathways involved in blistering disease?
The answer was not initially established.
But it created a plausible hypothesis that could be tested against clinical and epidemiological evidence.
4. Start With the Background Risk
Before evaluating a suspected drug-event association, the evaluator needs to understand how frequently the event occurs without exposure.
This is especially important for BP.
The disease predominantly affects older adults, and its incidence varies substantially between populations and study methods.
The patient population receiving DPP-4 inhibitors also tends to include older adults with multiple comorbidities.
Therefore, an observed increase in BP among DPP-4 inhibitor users could theoretically arise through several pathways:
- a true drug effect;
- confounding by age;
- confounding by diabetes;
- confounding by neurological disease;
- differences in healthcare utilisation;
- differences in diagnostic intensity;
- or a combination of these factors.
The evaluator should therefore avoid interpreting crude case counts without considering the background population.
5. The First Signal: Individual Cases
The early signal was driven in part by case reports describing BP occurring in patients exposed to DPP-4 inhibitors.
Individual cases can be highly informative when the event is:
- unusual;
- clinically well defined;
- temporally plausible;
- objectively confirmed;
- and associated with a credible dechallenge or other supportive evidence.
However, spontaneous reports have fundamental limitations.
They generally do not provide a denominator.
The pharmacovigilance team may know that 50 cases were reported.
It usually does not know directly from spontaneous reporting whether those 50 cases represent:
- 50 cases among 10,000 exposed patients;
- 50 cases among 1 million exposed patients;
- or 50 cases in a population with substantial background incidence.
This is why spontaneous reports are excellent for signal generation but usually insufficient for quantifying incidence or relative risk.
6. What Makes an Individual Case More Informative?
A case of BP occurring during DPP-4 inhibitor therapy becomes more informative when it contains:
- a clear diagnosis;
- appropriate dermatological investigations;
- a plausible time to onset;
- documentation of concomitant medicines;
- relevant medical history;
- information on age and comorbidities;
- treatment duration;
- dechallenge information;
- rechallenge information where available;
- and outcome.
For example:
An elderly patient develops biopsy-supported BP after prolonged exposure to a DPP-4 inhibitor, the medicine is discontinued, and the disease improves.
This is supportive evidence.
It is not proof.
The patient may have developed BP independently.
The evaluator should therefore use the case to ask whether the pattern is sufficiently unusual to justify examination of the broader evidence base.
7. Dechallenge: Useful but Not Definitive
Improvement after stopping a medicine can strengthen causality.
However, dechallenge is difficult to interpret for diseases such as BP.
The disease may fluctuate naturally.
Treatment for BP may also include:
- corticosteroids;
- immunosuppressants;
- biologic therapies;
- wound care;
- and other interventions.
Therefore:
Improvement after withdrawal does not automatically prove that the medicine caused the disease.
The strength of dechallenge evidence depends on:
- whether the suspected drug was the principal intervention changed;
- whether other treatments changed simultaneously;
- whether the timing of improvement is plausible;
- and whether the clinical course is consistent with the natural history of the disease.
This is an important general principle for signal evaluation.
8. Rechallenge
Rechallenge would theoretically provide strong evidence if the disease recurred following re-exposure.
However, intentional rechallenge with a medicine suspected of causing a serious autoimmune blistering disease would generally not be ethically appropriate merely to test causality.
Consequently, positive rechallenge information is likely to be rare.
The absence of rechallenge therefore should not be interpreted as evidence against causality.
This illustrates an important asymmetry in pharmacovigilance:
Some of the strongest causal evidence may be ethically unavailable.
The evaluator must therefore work with the evidence that can reasonably be obtained.
9. Disproportionality Analysis
As reports accumulated, pharmacovigilance database analysis became important.
Disproportionality analysis asks whether a particular drug-event combination is reported more frequently than expected relative to the background reporting experience of the database.
Common measures include:
- reporting odds ratio;
- proportional reporting ratio;
- information component;
- empirical Bayes measures.
A positive disproportionality result can indicate that a drug-event pair deserves attention.
But disproportionality is not incidence.
It is not relative risk.
And it is not a causal estimate.
A high disproportionality score can arise because:
- the event is genuinely associated with the medicine;
- the event has received unusual publicity;
- reporting is stimulated by regulatory attention;
- the medicine is used in a population at increased baseline risk;
- a reporting cluster exists;
- or other forms of reporting bias are present.
Therefore:
Disproportionality is a signal-detection tool, not a causal inference tool.
10. The Vildagliptin Signal
The European regulatory record provides a useful example of how the signal progressed.
In 2016, PRAC considered a signal for:
vildagliptin; vildagliptin/metformin — pemphigoid.
The PRAC documentation states that it considered available evidence from EudraVigilance and the literature and noted that the disproportionality score for vildagliptin was the highest among the DPP-4 inhibitor class.
PRAC recommended a variation to the product information.
The proposed wording added:
"Exfoliative and bullous skin lesions, including bullous pemphigoid"
with frequency listed as not known.
This is an excellent example of how a pharmacovigilance signal can move from database evidence into a concrete regulatory measure.
The important point is that the action did not require a precise estimate of absolute risk.
The regulator had sufficient evidence to justify communicating the potential risk.
[1][2]
11. What Does "Highest Disproportionality" Mean?
The phrase can easily be misunderstood.
Suppose one DPP-4 inhibitor has a higher reporting odds ratio than another.
That does not automatically mean:
The first drug is biologically more dangerous.
Possible explanations include:
- differences in utilisation;
- differences in duration of market exposure;
- stimulated reporting;
- reporting patterns in particular countries;
- different patient populations;
- differences in awareness;
- or genuine molecule-specific effects.
A pharmacovigilance evaluator should therefore treat class comparisons in spontaneous-reporting databases cautiously.
The disproportionality result tells us:
This drug-event combination is reported more disproportionately than expected.
It does not tell us:
This drug has the highest true incidence of the event.
12. Evidence From Case-Control Studies
Case-control studies subsequently provided more structured evidence.
One influential study examined patients with diabetes who developed BP and compared them with diabetic controls without BP.
The study found an association between DPP-4 inhibitor use and BP.
In the 2018 study by Kridin and Bergman, DPP-4 inhibitor use was associated with an approximately threefold increased risk of BP.
The association was particularly strong for vildagliptin and was also observed for linagliptin.
The study also found differences in clinical presentation between exposed and non-exposed patients with BP.
These findings were important because they moved the evidence beyond spontaneous reports.
However, case-control studies remain vulnerable to:
- selection bias;
- recall bias;
- confounding;
- exposure misclassification;
- and differences in diagnostic assessment.
Therefore, the results strengthened the signal without independently resolving causality.
[3]
13. Comparative Cohort Evidence
A major step in signal evaluation is to move toward comparative cohort designs.
A particularly informative study compared initiators of DPP-4 inhibitors with initiators of second-generation sulfonylureas.
The comparison is useful because it asks:
Among patients with type 2 diabetes starting treatment, is BP more common after starting a DPP-4 inhibitor than after starting another glucose-lowering treatment?
This is more informative than simply comparing DPP-4 inhibitor users with the general population.
The study included approximately 1.66 million patients across large US insurance and Medicare datasets.
After propensity-score matching, the pooled incidence rate was:
- 0.42 BP cases per 1,000 person-years among DPP-4 inhibitor initiators;
- 0.31 cases per 1,000 person-years among sulfonylurea initiators.
The pooled hazard ratio was approximately:
1.42 (95% CI 1.17–1.72).
The authors also reported higher relative risk in certain subgroups, including patients aged 65 years or older and white patients.
These findings provide both:
- an estimate of relative risk;
- and an estimate of absolute incidence.
That is a substantial improvement over spontaneous reporting alone.
[4]
14. Why the Absolute Risk Matters
Relative measures can sound dramatic when the underlying event is rare.
A hazard ratio around 1.4 indicates a relative increase.
But the absolute incidence remained low.
The study reported approximately:
0.42 cases per 1,000 person-years
among DPP-4 inhibitor initiators.
That corresponds to approximately:
42 cases per 100,000 person-years.
The comparator rate was approximately:
31 cases per 100,000 person-years.
The approximate absolute difference was therefore:
11 additional cases per 100,000 person-years, based on that study population and design.
This distinction is important.
A relative risk increase and an absolute risk increase answer different questions.
For pharmacovigilance decision-making, both matter.
15. Why the Comparator Matters
The choice of comparator can materially influence the result.
A comparison with the general population may be confounded by:
- diabetes;
- age;
- healthcare utilisation;
- comorbidities;
- and treatment indications.
An active comparator such as a sulfonylurea can reduce some of those differences.
It does not eliminate confounding.
Patients selected for one treatment may still differ systematically from patients selected for another.
Nevertheless, active-comparator new-user designs are generally much more informative for causal inference than crude exposed-versus-unexposed comparisons.
16. New-User Designs
A new-user design attempts to compare patients at a similar point in the treatment pathway.
This is important because prevalent users have already survived and tolerated treatment for some period.
If the analysis includes only prevalent users, early adverse events may be missed.
New-user designs therefore improve comparability around treatment initiation.
For a suspected drug-induced disease, this can help reduce certain forms of:
- survivor bias;
- exposure misclassification;
- and differences in treatment history.
The DPP-4 inhibitor/BP literature illustrates why study design should be examined alongside the numerical result.
17. Propensity Scores
Large observational studies increasingly use propensity scores to balance measured characteristics between treatment groups.
A propensity score estimates the probability of receiving one treatment based on observed covariates.
Matching or weighting can then make treatment groups more comparable on those measured characteristics.
This can improve causal inference.
But it does not solve unmeasured confounding.
For example, if a database does not adequately capture:
- frailty;
- neurological disease severity;
- smoking;
- dermatological history;
- socioeconomic factors;
- or other relevant variables,
propensity-score adjustment cannot fully correct those differences.
Therefore:
Propensity matching balances measured variables; it does not magically create randomisation.
18. Age as a Major Confounder
Age is particularly important in BP.
Older age increases baseline BP risk.
DPP-4 inhibitors are also frequently prescribed in older patients.
Therefore:
drug exposure + older age + BP
does not automatically demonstrate a drug effect.
A robust study must account for age.
The comparative cohort study found that the association was stronger among patients aged 65 years or older.
That finding is clinically interesting.
But it can be interpreted in more than one way.
It may indicate:
- increased susceptibility among older patients;
- a higher baseline BP rate producing greater absolute numbers;
- residual confounding;
- or a combination.
Subgroup findings therefore require careful interpretation.
19. Neurological Disease
BP has been associated with neurological disorders, including conditions such as dementia and cerebrovascular disease.
These conditions may also be common in older diabetic populations.
Therefore, neurological disease can act as a potential confounder.
This is a useful lesson because it demonstrates that the most important confounders are not necessarily obvious from the drug indication.
A signal evaluator should ask:
What characteristics are associated with the adverse event independently of the drug?
That question should be answered before interpreting the drug-event association.
20. Clinical Phenotype as Causal Evidence
An especially interesting feature of the DPP-4 inhibitor/BP signal is that exposed patients may show clinical differences from patients with apparently idiopathic BP.
Some studies reported differences in:
- mucosal involvement;
- eosinophil counts;
- autoantibody characteristics;
- and clinical course.
These observations are valuable.
A drug-associated phenotype that differs consistently from background disease can support a causal hypothesis.
But phenotype differences can also reflect:
- selection of cases;
- differences in disease severity;
- differences in diagnostic timing;
- or differences in treatment.
Therefore, phenotype should be integrated with epidemiology rather than treated as definitive proof.
21. Biological Plausibility
DPP-4 is not merely a glucose-metabolism enzyme.
It is involved in multiple biological pathways and is expressed on various cell types.
DPP-4/CD26 participates in:
- peptide metabolism;
- immune-cell signalling;
- T-cell biology;
- and interactions with cytokine and chemokine pathways.
This creates a biologically plausible basis for an immune-mediated adverse effect.
However, the exact mechanism by which DPP-4 inhibition might contribute to BP remains incompletely established.
This distinction should be explicit.
Established
DPP-4/CD26 has functions beyond glucose regulation and participates in immune biology.
Supported
DPP-4 inhibition has biological effects that could plausibly influence immune regulation.
Hypothesised
Those effects may contribute to the development of BP in susceptible patients.
Not established
A single definitive molecular pathway explaining all DPP-4 inhibitor-associated BP cases.
Mechanistic plausibility therefore strengthens the signal but does not substitute for epidemiological evidence.
22. Autoantibodies and BP180
BP is associated with autoantibodies against components of the basement membrane zone, especially BP180.
Some research has investigated whether DPP-4 inhibitor-associated BP has distinctive immunological characteristics.
These studies are valuable because they can potentially explain why exposure leads to disease in some patients but not others.
However, immunological findings need to be interpreted carefully.
A biomarker associated with exposure does not necessarily establish:
- temporal causation;
- specificity;
- or the mechanism responsible for disease initiation.
Mechanistic studies are most useful when they converge with epidemiological and clinical evidence.
23. Why a Class Effect Is Difficult to Establish
One of the most interesting questions is whether BP represents:
a DPP-4 inhibitor class effect
or
a risk concentrated in particular molecules.
The historical evidence initially drew particular attention to vildagliptin.
Subsequent studies reported associations with several DPP-4 inhibitors.
That pattern supports the possibility of a class-related biological mechanism.
However, the magnitude of association may differ between individual agents.
Potential explanations include:
- pharmacological differences;
- duration of exposure;
- prescribing populations;
- sample size;
- reporting behaviour;
- or chance.
A class effect should therefore be established through convergence of evidence, not by assuming that all drugs in a pharmacological class have identical risks.
24. The Importance of Consistency
Consistency is one of the most useful causal considerations.
The DPP-4 inhibitor/BP signal became more persuasive because evidence accumulated from different sources:
- spontaneous reports;
- disproportionality;
- case series;
- case-control studies;
- cohort studies;
- and comparative analyses.
The findings were not identical.
But the direction of evidence increasingly supported an association.
This is stronger than having one very large positive study.
At the same time, heterogeneity remained important.
Some datasets showed stronger associations than others.
The correct conclusion is therefore not:
Every study demonstrated the same risk.
Rather:
Multiple independent evidence streams supported an association, although the magnitude of the association varied between studies and individual active substances.
That is a more defensible statement.
25. A Systematic Review and Meta-Analysis
A systematic review and meta-analysis subsequently assessed medication associations with BP.
For DPP-4 inhibitors, the pooled case-control evidence showed an increased association, with a pooled odds ratio around 1.9 in the analysis.
The review also included cohort and randomised evidence.
Meta-analysis can improve statistical precision.
However, pooling observational studies does not eliminate their biases.
A pooled estimate may still inherit:
- confounding;
- heterogeneity;
- exposure misclassification;
- outcome misclassification;
- publication bias;
- and differences in study design.
Therefore:
Meta-analysis increases precision only to the extent that the underlying studies are sufficiently comparable and valid.
A narrow confidence interval around a biased estimate does not make the estimate causal.
[5]
26. Heterogeneity Is Information
Suppose one study reports:
OR 3.2
another:
HR 2.2
and another:
HR 1.4.
It would be tempting to conclude that the "true risk" lies somewhere around 2.
That may be inappropriate.
The studies may differ in:
- population;
- comparator;
- age;
- diagnostic criteria;
- exposure duration;
- individual DPP-4 inhibitor;
- database;
- adjustment;
- and outcome ascertainment.
Instead of asking:
What is the average number?
the evaluator should first ask:
Why are the estimates different?
Heterogeneity can reveal important features of the signal.
27. Absolute Risk Versus Relative Risk
The clinical significance of the signal depends on both relative and absolute risk.
For example, a relative increase of approximately 40% may sound substantial.
But if the baseline event rate is very low, the absolute increase may remain small.
The approximate numbers from the large comparative cohort illustrate this.
| Measure | DPP-4 inhibitor | Sulfonylurea |
|---|---|---|
| Incidence per 1,000 person-years | 0.42 | 0.31 |
| Incidence per 100,000 person-years | 42 | 31 |
| Pooled HR | 1.42 | Reference |
| Approximate absolute difference | 11/100,000 person-years | Reference |
These figures belong to the specific study population and should not be treated as universal incidence estimates.
They demonstrate a broader point:
Regulatory significance depends on more than the relative effect estimate.
Severity, preventability, susceptible populations and available alternatives also matter.
28. Regulatory Interpretation
The European regulatory record provides a particularly useful example of proportionate action.
In 2016, PRAC considered the vildagliptin/pemphigoid signal and recommended an amendment to product information.
The proposed SmPC wording placed:
"Exfoliative and bullous skin lesions, including bullous pemphigoid"
under skin and subcutaneous tissue disorders, with frequency listed as not known.
This is important because the action was not equivalent to declaring that all patients treated with the medicine would develop BP.
Instead, the regulatory action communicated the identified potential adverse reaction.
[1][2]
29. Why "Frequency Not Known" Is Meaningful
A frequency category of "not known" does not mean:
the event is extremely rare.
It means the available data are insufficient to make a reliable frequency estimate under the applicable product-information framework.
This distinction is important.
A pharmacovigilance evaluator should not interpret:
Frequency not known
as:
no meaningful evidence exists.
The signal may be sufficiently credible to warrant inclusion while still lacking a robust denominator.
This is another example of the difference between:
- signal recognition;
- frequency estimation;
- and causal certainty.
30. What the Regulatory Action Did Not Establish
The product-information change did not establish:
- the exact incidence of BP;
- that every DPP-4 inhibitor carries identical risk;
- that every case occurring during treatment is drug-induced;
- the precise biological mechanism;
- or the exact magnitude of the causal effect.
It established something narrower and clinically useful:
There was sufficient evidence to communicate BP as a potential adverse reaction.
That distinction should be maintained when describing historical regulatory decisions.
31. What a Signal Evaluator Should Ask at Each Stage
The case can be converted into a practical evaluation sequence.
Stage 1 — Detection
Ask:
- Is the event unusual?
- Is the diagnosis reasonably specific?
- Is there more reporting than expected?
- Is there a plausible temporal relationship?
Stage 2 — Validation
Ask:
- Are the cases medically credible?
- Are alternative explanations present?
- Is the event serious?
- Is the association already known?
Stage 3 — Characterisation
Ask:
- Which active substances are involved?
- Is there a class pattern?
- What is the latency?
- What is the clinical phenotype?
- Is there dechallenge?
Stage 4 — Epidemiological assessment
Ask:
- What is the background incidence?
- What comparator is appropriate?
- Is the association consistent?
- Is there dose or duration response?
- Have major confounders been addressed?
Stage 5 — Mechanistic assessment
Ask:
- Is there biological plausibility?
- Is there a plausible immune pathway?
- Is there evidence of a drug-specific phenotype?
- What remains unknown?
Stage 6 — Regulatory assessment
Ask:
- Is the event clinically serious?
- Is the evidence sufficiently credible?
- Is there a meaningful preventable risk?
- Can product information or other measures reduce risk?
32. A Worked Individual Case
Consider a hypothetical patient:
A 76-year-old man with type 2 diabetes begins a DPP-4 inhibitor.
After 14 months, he develops:
- severe pruritus;
- tense blisters on the trunk and limbs;
- erosions;
- and progressive skin involvement.
Dermatology assessment demonstrates findings compatible with BP.
A biopsy and direct immunofluorescence support the diagnosis.
The DPP-4 inhibitor is discontinued.
The patient receives appropriate treatment for BP and subsequently improves.
How should the case be evaluated?
Evidence supporting causality
- medically well-defined event;
- plausible temporal relationship;
- objective diagnostic evidence;
- biologically plausible class;
- known pharmacovigilance signal;
- supportive epidemiological evidence;
- improvement after withdrawal.
Evidence against or limiting certainty
- advanced age;
- diabetes itself;
- possible neurological comorbidities;
- spontaneous background incidence of BP;
- absence of rechallenge;
- potential contribution from other medicines.
The correct conclusion is not:
Definite drug-induced BP.
Nor is it:
Unrelated because the patient is elderly.
The appropriate conclusion depends on the complete case and the applicable causality framework.
33. Why the Individual Case and Population Signal Must Be Connected
The individual case provides detail.
The epidemiological study provides population-level context.
The spontaneous-reporting database identifies patterns.
The mechanistic literature provides biological context.
The regulator integrates the evidence.
These evidence streams answer different questions.
| Evidence source | Main contribution |
|---|---|
| Individual case | Clinical phenotype and chronology |
| Case series | Pattern recognition |
| Disproportionality | Signal detection |
| Case-control study | Association estimate |
| Cohort study | Incidence and comparative risk |
| Mechanistic research | Biological plausibility |
| Meta-analysis | Overall evidence synthesis |
| Regulatory assessment | Clinical and public-health significance |
A high-quality signal evaluation should therefore avoid allowing one evidence source to dominate simply because it appears numerically impressive.
34. Common Analytical Errors
Error 1 — Treating disproportionality as incidence
A reporting odds ratio is not an incidence rate.
Error 2 — Treating a case report as proof
A compelling case can generate a signal but cannot establish population-level causality.
Error 3 — Ignoring background disease risk
BP occurs without DPP-4 inhibitor exposure.
Error 4 — Assuming class effect automatically
Evidence for one molecule does not automatically establish identical risk for all molecules.
Error 5 — Treating a positive epidemiological result as definitive
Observational studies remain vulnerable to residual confounding.
Error 6 — Treating a negative study as definitive reassurance
Rare outcomes can leave substantial uncertainty.
Error 7 — Ignoring absolute risk
Relative risk can exaggerate the perceived clinical magnitude of a rare event.
Error 8 — Overstating mechanism
A plausible immune pathway is not the same as an established mechanism.
Error 9 — Ignoring phenotype
Differences in clinical presentation may provide important mechanistic clues.
Error 10 — Assuming regulatory wording equals causal certainty
Product information reflects a regulatory risk-management decision, not necessarily complete scientific resolution.
35. What Is Established, What Is Supported and What Remains Uncertain?
A useful way to communicate the evidence is to separate different levels of certainty.
Established
- Bullous pemphigoid is a recognised autoimmune blistering disorder.
- DPP-4 inhibitors are widely used glucose-lowering medicines.
- Regulatory authorities identified and evaluated a signal involving DPP-4 inhibitors and pemphigoid.
- EMA's PRAC recommended product-information changes for vildagliptin-containing products in 2016.
- Epidemiological studies have reported an association between DPP-4 inhibitor exposure and BP.
Strongly supported
- The association is not adequately explained by a single isolated case.
- The signal is supported by multiple evidence streams.
- The association appears clinically relevant enough to warrant awareness and risk management.
Plausible but incompletely established
- DPP-4/CD26-related immune effects may contribute to susceptibility.
- Some patients may have a distinct drug-associated BP phenotype.
- Risk may differ between individual DPP-4 inhibitors.
Still uncertain
- The exact causal mechanism.
- The precise absolute excess risk for individual active substances.
- The degree to which the observed association varies by age, sex, ethnicity and other susceptibility factors.
- Whether the magnitude of risk is identical across the entire DPP-4 inhibitor class.
- Which patients are biologically most susceptible.
This structure prevents the article from converting hypotheses into facts.
36. Regulatory Significance Versus Causal Certainty
One of the most important lessons from this signal is that the regulatory threshold and the scientific threshold are not identical.
A regulator may conclude:
The evidence is sufficient to warrant a warning.
That does not necessarily mean:
The exact causal mechanism and absolute risk have been established.
The two statements can coexist.
Pharmacovigilance is fundamentally concerned with decision-making under uncertainty.
The relevant question is often:
What action is proportionate to the available evidence and the potential consequence of failing to act?
37. Risk Minimisation
For a rare but potentially serious adverse event such as BP, risk minimisation may include:
- awareness among prescribers;
- recognition of blistering and pruritus;
- prompt dermatological assessment;
- consideration of discontinuing the suspected medicine when BP is diagnosed;
- appropriate product-information wording;
- and continued pharmacovigilance surveillance.
Risk minimisation should be proportionate.
A signal does not automatically require withdrawal of a medicine.
The decision depends on:
- seriousness;
- frequency;
- strength of evidence;
- availability of alternatives;
- reversibility;
- patient susceptibility;
- and the overall benefit-risk balance.
38. Why the Signal Matters for Aggregate Safety Review
The DPP-4 inhibitor/BP signal also demonstrates why aggregate review cannot be reduced to counting cases.
A useful aggregate assessment should examine:
- case volume;
- reporting trends;
- disproportionality;
- seriousness;
- diagnostic confirmation;
- time to onset;
- dechallenge;
- rechallenge;
- age distribution;
- sex distribution;
- individual active substance;
- concomitant medications;
- neurological disease;
- clinical phenotype;
- epidemiological evidence;
- and regulatory developments.
The question is not:
How many cases occurred?
It is:
Does the accumulating evidence change the interpretation of the drug-event relationship or the adequacy of existing risk minimisation?
39. A Practical Evidence Table for Signal Review
A working signal-evaluation table might look like this:
| Domain | Question | Finding | Interpretation |
|---|---|---|---|
| Cases | Are medically credible cases present? | Yes | Supports signal |
| Temporal relationship | Is onset compatible? | Variable but plausible | Supportive |
| Dechallenge | Does disease improve after withdrawal? | Reported in cases | Supportive but not definitive |
| Rechallenge | Does disease recur? | Rare/unavailable | Limited evidence |
| Disproportionality | Is reporting disproportionate? | Yes, particularly for vildagliptin | Signal detection support |
| Case-control studies | Is there an association? | Yes in several studies | Supports association |
| Cohort studies | Is comparative risk increased? | Yes in large studies | Strengthens evidence |
| Absolute risk | Is the event common? | No | Important for benefit-risk |
| Confounding | Are major confounders present? | Yes | Limits causal certainty |
| Biological plausibility | Is there a credible mechanism? | Plausible | Supportive |
| Class consistency | Are several molecules implicated? | Yes | Supports class consideration |
| Regulatory action | Has product information changed? | Yes for vildagliptin | Regulatory significance |
| Overall conclusion | Is causality certain? | No | Residual uncertainty remains |
40. Lessons for Signal Management
Lesson 1 — Start with a precise case definition
A nonspecific "rash" signal cannot be evaluated in the same way as biopsy-supported BP.
Lesson 2 — Separate signal detection from signal assessment
A disproportionality result tells you what deserves investigation.
It does not tell you what caused the event.
Lesson 3 — Use the right denominator
Spontaneous reports generally cannot provide incidence.
Epidemiological studies can.
Lesson 4 — Use comparative designs where possible
Active-comparator and new-user designs can improve causal inference.
Lesson 5 — Examine absolute as well as relative risk
A serious event can be clinically important even when rare, but relative risk alone does not describe the patient-level magnitude.
Lesson 6 — Investigate confounding explicitly
Age, diabetes and neurological disease are not minor footnotes in BP evaluation.
Lesson 7 — Examine class versus molecule-specific patterns
Do not assume either a universal class effect or a molecule-specific effect without evidence.
Lesson 8 — Use phenotype as supporting evidence
Distinct clinical or immunological features may strengthen a causal hypothesis.
Lesson 9 — Explain mechanisms carefully
Separate established biology from plausible hypotheses.
Lesson 10 — Do not wait for perfect certainty before considering risk minimisation
Regulatory action can be proportionate to credible potential risk.
Lesson 11 — Treat regulatory wording as evidence of regulatory significance
It is not necessarily proof of a quantified causal effect.
Lesson 12 — Keep reassessing the signal
The interpretation of a signal can change as epidemiological and mechanistic evidence accumulates.
41. What This Case Teaches a QPPV
A QPPV reviewing this signal should be able to answer several questions without relying solely on the latest aggregate case count.
Is the event medically credible?
Yes.
BP is a well-defined clinical disease.
Is the association biologically plausible?
Yes, but the precise mechanism remains incompletely established.
Is the signal supported beyond spontaneous reports?
Yes.
Case-control and cohort studies have provided epidemiological support.
Is the risk precisely quantified?
No.
The magnitude varies between studies and may differ between active substances.
Are there important alternative explanations?
Yes.
Age, diabetes, neurological disease and other patient characteristics require consideration.
Is there regulatory significance?
Yes.
PRAC recommended product-information changes for vildagliptin-containing products in 2016.
Does a positive signal mean every case is caused by the medicine?
No.
The background incidence and patient-specific risk remain important.
Does uncertainty mean the signal should be ignored?
No.
The regulatory response demonstrates why clinically meaningful uncertainty can still require action.
42. Overall Signal Assessment
The accumulated evidence supports an association between DPP-4 inhibitor exposure and bullous pemphigoid.
The signal was initially strengthened by:
- medically credible case reports;
- spontaneous-reporting patterns;
- and disproportionality, particularly for vildagliptin.
It was subsequently supported by:
- case-control studies;
- cohort studies;
- comparative pharmacoepidemiology;
- and systematic reviews.
The association is biologically plausible because DPP-4/CD26 has functions in immune regulation, although the precise mechanism linking pharmacological DPP-4 inhibition to BP remains incompletely established.
The epidemiological evidence is not perfectly uniform.
The magnitude of association varies between studies and active substances.
Important confounding factors remain, particularly because BP is strongly associated with older age and because patients receiving DPP-4 inhibitors may have characteristics that differ from comparator populations.
The evidence therefore supports a clinically meaningful safety signal but does not justify claiming that every case of BP occurring during DPP-4 inhibitor therapy is caused by the medicine.
The most defensible overall conclusion is:
Multiple independent evidence streams support an association between DPP-4 inhibitor exposure and bullous pemphigoid. The signal progressed from spontaneous reports and disproportionality to epidemiological evidence demonstrating increased relative risk in several comparative studies. The association is biologically plausible, but the exact mechanism, absolute excess risk and degree of molecule-specific susceptibility remain incompletely established. Regulatory action, including product-information changes for vildagliptin-containing products, was therefore proportionate to a credible and clinically relevant potential risk rather than dependent on complete mechanistic certainty.
43. Final Perspective
The DPP-4 inhibitor and bullous pemphigoid signal is a good example of why signal evaluation should be understood as an evidence-building process.
The first case does not answer the question.
The disproportionality analysis does not answer the question.
The first epidemiological study does not answer the question.
The meta-analysis does not necessarily answer the question.
Each adds another piece of evidence.
The evaluator's task is to determine whether the pieces converge.
In this case, they increasingly did.
The signal moved from:
unexpected cases
to:
disproportionate reporting
to:
epidemiological association
to:
regulatory recognition
while uncertainty remained around:
- exact mechanism;
- absolute risk;
- class versus molecule-specific effects;
- and individual susceptibility.
That is not a weakness of the assessment.
It is the reality of pharmacovigilance.
A strong signal evaluation does not hide uncertainty.
It identifies it, explains why it exists, and determines whether the remaining uncertainty is compatible with continued use, additional investigation, or risk-minimisation measures.
For QPPVs and signal evaluators, this is the central lesson:
The quality of a signal assessment is determined not by how confidently it states a conclusion, but by how rigorously it connects the evidence to the conclusion and makes the remaining uncertainty visible.
References
-
European Medicines Agency. PRAC recommendations on signals adopted at the PRAC meeting of 28 November–1 December 2016. EMA/PRAC/740369/2016. Vildagliptin; vildagliptin/metformin — pemphigoid (EPITT No. 18692).
-
European Medicines Agency. New product information wording: extracts from PRAC recommendations on signals adopted at the 28 November–1 December 2016 PRAC meeting. EMA/PRAC/740435/2016.
-
Kridin K, Bergman R. Association of Bullous Pemphigoid With Dipeptidyl-Peptidase 4 Inhibitors in Patients With Diabetes: Estimating the Risk of the New Agents and Characterizing the Patients. JAMA Dermatology. 2018;154(10):1152–1158. doi:10.1001/jamadermatol.2018.2352.
-
Lee H, Chung HJ, Pawar A, Patorno E, Kim DH. Evaluation of Risk of Bullous Pemphigoid With Initiation of Dipeptidyl Peptidase–4 Inhibitor vs Second-generation Sulfonylurea. JAMA Dermatology. 2020;156(5):508–515. doi:10.1001/jamadermatol.2020.0267.
-
Phan K, Charlton O, Smith SD. A systematic review and meta-analysis of the association between medication use and bullous pemphigoid. JAMA Dermatology. 2020;156(7):742–752.
-
Lee SG, Lee HJ, Yoon MS, Kim DH. Association of Dipeptidyl Peptidase 4 Inhibitor Use With Risk of Bullous Pemphigoid in Patients With Diabetes. JAMA Dermatology. 2019;155(2):172–177. doi:10.1001/jamadermatol.2018.4556.
-
GarcĂa-DĂez I, Ivars-LleĂł M, LĂłpez-AventĂn D, et al. Bullous pemphigoid induced by dipeptidyl peptidase-4 inhibitors: eight cases with clinical and immunological characterization. International Journal of Dermatology. 2018;57(7):810–816.
-
Yoshiji S, Murakami T, Harashima SI, et al. Bullous pemphigoid associated with dipeptidyl peptidase-4 inhibitors: a report of five cases. Journal of Diabetes Investigation. 2018;9(2):445–447.
-
European Medicines Agency. Guideline on good pharmacovigilance practices (GVP) Module IX — Signal Management (Rev. 1).
-
European Medicines Agency. Questions and answers on signal management.
-
European Medicines Agency. Guideline on good pharmacovigilance practices (GVP) Module V — Risk Management Systems.
-
European Medicines Agency. Guideline on good pharmacovigilance practices (GVP) Module XVI — Risk Minimisation Measures: Selection of Tools and Effectiveness Indicators.
-
European Commission. Commission Implementing Regulation (EU) No 520/2012 on the performance of pharmacovigilance activities.