Statistics Interpretation

P-Values and Statistical Significance: What They Do Not Prove

A p-value is one output from a hypothesis test, not a verdict on importance or truth. Read it alongside the study design, effect estimate, uncertainty, and number of tests performed.

Published by PeptideSchool Editorial DeskPublished 2026-08-11Reviewed 2026-08-11

Educational content only. Not medical advice.

A p-value is conditional on a model and null hypothesis

A p-value summarizes how incompatible the observed data, or more extreme data defined by the test, are with a specified statistical model that includes the null hypothesis. It is not the probability that the null hypothesis is true, not the probability that results occurred 'by chance,' and not the chance that a replication will succeed. Interpretation depends on the test, sampling process, analysis plan, and assumptions that produced it.

A threshold is a decision convention, not a truth boundary

The familiar 0.05 threshold does not create a sharp scientific difference between p = 0.049 and p = 0.051. Labeling one result significant and the other not significant can conceal nearly identical evidence. Exact values, effect estimates, confidence intervals, outcome importance, and prior evidence provide a fuller account. A small p-value can accompany a trivial effect in a large sample, while a meaningful effect can remain uncertain in a small sample.

Multiplicity and analytical flexibility change the error landscape

Testing many outcomes, subgroups, time points, or models increases opportunities for small p-values even without a stable underlying effect. Pre-specification, correction methods, transparent reporting, and independent replication help readers evaluate that risk. Selectively highlighting the smallest p-value while omitting the total number of analyses gives an incomplete picture. Post-hoc results can generate hypotheses but should not be presented as if they were the sole confirmatory test.

Read p-values as one component of an evidence chain

Begin with the research question and design, then inspect the primary outcome, effect magnitude, interval precision, missing data, protocol deviations, and biological plausibility. Ask whether the analysis was planned and whether results persist across sensitivity analyses or independent studies. Statistical significance does not establish causality, practical importance, clinical relevance, or applicability to a particular person. Those are separate judgments requiring more evidence.

Evidence limits

  • Different statistical frameworks define and use evidence differently.
  • A public article cannot verify assumptions or undisclosed analyses in a specific study.
  • Statistical evidence does not translate directly into individualized medical relevance.

Sources and further reading

These sources ground the definitions and evidence boundaries on this page. A citation is a route for verification, not an endorsement of a product or personal use.

American Statistical Association

The ASA's Statement on p-Values: Context, Process, and Purpose

Primary consensus statement describing principles and common p-value misinterpretations.

Open source

Jornal Vascular Brasileiro via PubMed Central

P-value and Effect-size in Clinical and Experimental Studies

Peer-reviewed discussion of interpreting p-values with effect sizes and confidence intervals.

Open source

Common questions

Does p < 0.05 prove the hypothesis is true?

No. It is a model-conditional data summary and does not assign a probability to the truth of a hypothesis.

Does p > 0.05 prove there is no effect?

No. The study may be imprecise, underpowered, or compatible with effects in several directions.

Why report effect sizes with p-values?

Effect sizes describe magnitude, while confidence intervals describe precision; a p-value alone provides neither.

Continue with context

Build a more rigorous reading workflow

Explore the complete PeptideSchool research workspace, including source trails, study organization, and the Premium Calculator. The tools are educational and are not a substitute for professional judgment.

Explore Premium