Statistics Interpretation
P-Values and Statistical Significance: What They Do Not Prove
A p-value is one output from a hypothesis test, not a verdict on importance or truth. Read it alongside the study design, effect estimate, uncertainty, and number of tests performed.
Educational content only. Not medical advice.
A p-value is conditional on a model and null hypothesis
A p-value summarizes how incompatible the observed data, or more extreme data defined by the test, are with a specified statistical model that includes the null hypothesis. It is not the probability that the null hypothesis is true, not the probability that results occurred 'by chance,' and not the chance that a replication will succeed. Interpretation depends on the test, sampling process, analysis plan, and assumptions that produced it.
A threshold is a decision convention, not a truth boundary
The familiar 0.05 threshold does not create a sharp scientific difference between p = 0.049 and p = 0.051. Labeling one result significant and the other not significant can conceal nearly identical evidence. Exact values, effect estimates, confidence intervals, outcome importance, and prior evidence provide a fuller account. A small p-value can accompany a trivial effect in a large sample, while a meaningful effect can remain uncertain in a small sample.
Multiplicity and analytical flexibility change the error landscape
Testing many outcomes, subgroups, time points, or models increases opportunities for small p-values even without a stable underlying effect. Pre-specification, correction methods, transparent reporting, and independent replication help readers evaluate that risk. Selectively highlighting the smallest p-value while omitting the total number of analyses gives an incomplete picture. Post-hoc results can generate hypotheses but should not be presented as if they were the sole confirmatory test.
Read p-values as one component of an evidence chain
Begin with the research question and design, then inspect the primary outcome, effect magnitude, interval precision, missing data, protocol deviations, and biological plausibility. Ask whether the analysis was planned and whether results persist across sensitivity analyses or independent studies. Statistical significance does not establish causality, practical importance, clinical relevance, or applicability to a particular person. Those are separate judgments requiring more evidence.
Evidence limits
- Different statistical frameworks define and use evidence differently.
- A public article cannot verify assumptions or undisclosed analyses in a specific study.
- Statistical evidence does not translate directly into individualized medical relevance.
Sources and further reading
These sources ground the definitions and evidence boundaries on this page. A citation is a route for verification, not an endorsement of a product or personal use.
American Statistical Association
The ASA's Statement on p-Values: Context, Process, and Purpose
Primary consensus statement describing principles and common p-value misinterpretations.
Open sourceJornal Vascular Brasileiro via PubMed Central
P-value and Effect-size in Clinical and Experimental Studies
Peer-reviewed discussion of interpreting p-values with effect sizes and confidence intervals.
Open sourceCommon questions
Does p < 0.05 prove the hypothesis is true?
No. It is a model-conditional data summary and does not assign a probability to the truth of a hypothesis.
Does p > 0.05 prove there is no effect?
No. The study may be imprecise, underpowered, or compatible with effects in several directions.
Why report effect sizes with p-values?
Effect sizes describe magnitude, while confidence intervals describe precision; a p-value alone provides neither.
Continue with context
Build a more rigorous reading workflow
Explore the complete PeptideSchool research workspace, including source trails, study organization, and the Premium Calculator. The tools are educational and are not a substitute for professional judgment.