Statistical interpretation
Multiple comparisons: why more tests create more chances for false positives
Every hypothesis test has a chance of a false-positive conclusion. When many related tests are searched without a prespecified strategy, the probability that at least one looks positive by chance can increase substantially.
Educational content only. Not medical advice.
Multiplicity appears in more places than endpoint lists
Trials can test several endpoints, treatment groups, time points, populations, models, subgroup cuts, or interim looks. Each added opportunity to declare a favorable result can increase the chance of at least one false-positive conclusion. The relevant family of tests depends on the decision being made, so counting p-values after publication is not a complete multiplicity assessment.
Error control begins with hierarchy and prespecification
A protocol can define one primary endpoint, co-primary requirements, a fixed testing sequence, gatekeeping families, or another justified adjustment. These plans allocate the tolerated Type I error before results are known. Methods differ in assumptions and power, but all aim to prevent a flexible search from being presented as one clean confirmatory test.
Exploration remains legitimate when labeled
Not every exploratory analysis needs to be suppressed. Broad examination can discover signals, interactions, or future endpoints. The protection is transparent labeling, complete reporting, effect estimates with uncertainty, and independent confirmation. A nominal p-value from a large unadjusted search should be treated as a hypothesis, not as isolated proof.
Read the whole analysis family
Ask how many outcomes, groups, time points, and subgroups were analyzed; which were prespecified; what adjustment or hierarchy was used; and whether the published claim follows that plan. Selective reporting can hide the denominator of unsuccessful analyses. Registries and protocols help reconstruct the family that the final paper alone may not display.
Evidence limits
- No single adjustment is optimal for every scientific objective or correlation structure.
- Multiplicity control does not make an endpoint valid or clinically meaningful.
- Exploratory analyses can remain informative when their status and uncertainty are explicit.
Sources and further reading
These sources ground the definitions and evidence boundaries on this page. A citation is a route for verification, not an endorsement of a product or personal use.
U.S. Food and Drug Administration
Multiple Endpoints in Clinical Trials
Official guidance on endpoint families, multiplicity, prespecification, and Type I error control.
Open sourceU.S. Food and Drug Administration / ICH
E9(R1) Statistical Principles for Clinical Trials: Estimands and Sensitivity Analysis
Official framework connecting trial objectives, estimands, intercurrent events, analysis, and interpretation.
Open sourceNational Institutes of Health
Clinical Trial Reporting Requirements
Official NIH policy context for registration, results reporting, and transparency.
Open sourceCommon questions
Does testing ten outcomes guarantee one false positive?
No. It raises the probability, with the exact risk depending on thresholds and correlation among tests.
Is Bonferroni the only correction?
No. Hierarchical, gatekeeping, stepwise, and other methods may be appropriate for different designs.
Can an unadjusted subgroup result be useful?
Yes as an explicitly exploratory signal, but it normally needs confirmation in a new, prespecified analysis.
Continue with context
Keep building your evidence-reading skills
Explore the public PeptideSchool research library for more source-backed methods, glossaries, and evidence maps.