PsyMeasureFDR

PsyMeasure · Statistics lab

False Discovery Rate (FDR) Simulator

Compare Bonferroni, Holm and Benjamini–Hochberg in a free multiple-testing simulator. Explore false discovery rate (FDR), FWER and statistical power with 11 methods.

How many significant results are just noise? Explore the trade-off between false discoveries and missed effects, using the same simulated experiments for every method.

Compare the outcomes

Ready to explore

See the simulation

Each tile is one test in a simulated experiment. Colour shows the truth; the mark shows what the method decided.

Results will appear here after the simulation.

Colour = true effect:negativenonepositive
Brightness:how strong the effect was in this sample (|z|)
Tile order:

The same experiment, method by method

Many experiments at a glance

Each row is one whole experiment: real effects on the left, noise on the right. Bright tiles on the left are found effects; yellow tiles on the right are false discoveries. Select a row to inspect it above.

False discovery rate (FDR)

Average false share among discoveries. Experiments with no discoveries contribute zero.

Statistical power

Share of real effects detected. Higher means fewer real effects missed.

Family-wise error (FWER)

Share of experiments containing at least one false discovery.

The target applies to FDR or FWER according to the method, not to power. Simulation estimates fluctuate around theoretical values.

Exact results

Averages across completed experiments. FP = false discoveries; FN = missed real effects.

Exact results
MethodFDRPowerFWERMean FPMean FN
Where the p-values come from

Null p-values are approximately uniform. Real effects tend toward zero. Each group is normalized separately; empty groups have no bars.

Null hypotheses (noise)Real effects

What this simulation assumes

Each experiment contains independent, two-sided, one-sample z-tests with known variance. Null tests use Z ~ N(0, 1); real effects use Z ~ N(d√n, 1). The number of real effects is rounded to a whole number. This model does not simulate correlated tests, t-tests, optional stopping or publication bias.

These are Monte Carlo estimates, not guarantees for one experiment. With 2,000 repetitions, an estimated FWER of 5% has a standard error of about 0.5 percentage points. FDR is the mean of experiment-level false discovery proportions, not the ratio of pooled false and total discoveries.

FDR / FWER

Multiple comparisons: FDR vs FWER

Testing many hypotheses creates more opportunities for false positives. Corrections change the threshold for calling a result significant. The right comparison is both error control and the ability to detect real effects.

Why multiple testing needs a correction

With 20 independent true null hypotheses and α = 0.05, the chance of at least one false positive is 1 − (1 − 0.05)²⁰ ≈ 64%. The expected number of false positives is 20 × 0.05 = 1. These are different quantities: a probability and a count.

What is the false discovery rate?

FDR = E[V / max(R, 1)], where V is the number of false discoveries and R is the total number of discoveries. A 5% FDR target controls this proportion on average across repeated experiments under the method's assumptions. It does not mean every significant result has a 5% probability of being false.

How is FWER different?

FWER = P(V ≥ 1): the probability of making even one false discovery in a family of tests. Bonferroni and Holm target FWER; Benjamini–Hochberg targets FDR. Controlling FWER also bounds FDR, but often detects fewer real effects.

Worked example: Bonferroni vs BH

Take five sorted p-values: 0.001, 0.008, 0.039, 0.041, 0.300, with α = q = 0.05. Bonferroni uses 0.05 / 5 = 0.01 and rejects the first two. BH compares ranks with 0.01, 0.02, 0.03, 0.04, 0.05. The largest passing rank is 2, so BH also rejects the first two. If the last p-value were 0.045, BH would reject all five: use the largest passing rank, not the first failing one.

Compare all 11 correction methods

Uncorrected

Reject each hypothesis when p < α. No family-level error control. This is a useful baseline for seeing how false positives accumulate.

Bonferroni

Bonferroni rejects when p < α/m. It controls FWER for valid p-values under any dependence structure, but can miss real effects when many tests are performed.

Holm

Sort p-values from smallest to largest. Compare rank r with α/(m−r+1), stopping at the first failure. Holm controls FWER under arbitrary dependence and rejects at least everything Bonferroni rejects.

Šidák

Use the single threshold 1−(1−α)^(1/m). Šidák controls FWER for independent tests and is slightly less conservative than Bonferroni.

Holm–Šidák

Apply the step-down logic of Holm with thresholds 1−(1−α)^(1/(m−r+1)). This simulation uses independent tests; do not assume this method handles arbitrary dependence.

Hochberg

Find the largest rank r satisfying p(r) ≤ α/(m−r+1), and reject all ranks through r. Hochberg controls FWER under independence or appropriate positive dependence conditions.

Hommel

Hommel uses a closed-testing procedure based on Simes tests. It controls FWER under independence or conditions supporting the Simes inequality, and is at least as powerful as Hochberg. It is more computationally expensive.

Benjamini–Hochberg

Find the largest rank r satisfying p(r) ≤ (r/m)q, then reject all ranks through r. Benjamini–Hochberg controls FDR under independent tests and certain positive dependence conditions (PRDS).

Benjamini–Yekutieli

Use BH with q divided by Hm = Σ(1/k), k = 1…m. Benjamini–Yekutieli controls FDR under arbitrary dependence, usually at a cost in power.

Storey-style (exploratory)

This simplified exploratory variant estimates π₀ from the fraction of p-values above 0.5, then runs BH at q/π̂₀. It is not a full q-value implementation and does not provide a universal finite-sample FDR guarantee.

Adaptive BH (exploratory)

This exploratory two-pass variant first runs BH at q, estimates m₀ as max(1, m−R₁), then runs BH at q·m/m̂₀. It is not the BKY procedure and is not presented as having its finite-sample guarantee.

Common questions

Is this an FDR calculator for my own p-values?

This is an educational simulator, not a p-value upload calculator. It generates experiments with known ground truth so you can observe false discoveries and missed effects, which cannot usually be known from real data alone.

Does a q-value of 0.05 guarantee only 5% false positives?

No. A target q = 0.05 concerns an expected proportion among discoveries under the procedure's assumptions. A single experiment can have a larger or smaller false discovery proportion.

Which correction should I choose?

Decide which errors your research design must control. FWER methods limit the risk of any false discovery in the test family. FDR methods control the average false share among discoveries. Check dependence assumptions and define the hypothesis family before examining the results.

Sources and further reading