Research Skill Center All articles
Research Quality & Integrity

The Number That Misleads: Moving Past P-Values to Communicate Research That Actually Matters

Research Skill Center
The Number That Misleads: Moving Past P-Values to Communicate Research That Actually Matters

For many researchers trained in the United States, the moment of truth in any quantitative study arrives when the analysis spits out a p-value. If it falls below 0.05, the study is a success. If it does not, something has gone wrong. This binary framework is so deeply embedded in graduate training, peer review culture, and journal expectations that it can feel less like a convention and more like a natural law.

It is neither. And its dominance has contributed to one of the most serious credibility challenges the scientific community has faced in a generation.

What a P-Value Actually Tells You

Before examining the limitations of p-values, it is worth being precise about what they do and do not communicate. A p-value is the probability of observing a result at least as extreme as the one you obtained, assuming that the null hypothesis is true. That is a mouthful, and it is notably not the same as the probability that your hypothesis is correct, or the probability that your result occurred by chance, or the magnitude of the effect you observed.

Yet in practice, p-values are routinely interpreted in all three of those ways—by students, by journalists, and sometimes by researchers who know better but fall back on familiar shorthand. The American Statistical Association issued a formal statement in 2016 explicitly cautioning against binary significance thresholds and the misinterpretation of p-values, and followed it with an expanded discussion in 2019. The message from the professional statistical community has been consistent: the p-value was never designed to carry the interpretive weight we have placed on it.

Large Samples and the Significance Illusion

One of the clearest demonstrations of the p-value's limitations is what happens when sample sizes become very large. In a study with tens of thousands of participants, even a trivially small difference between groups can produce a p-value well below 0.001. The result is technically significant by any conventional threshold, but the practical meaning may be negligible.

Imagine a study examining the effect of a workplace wellness app on employee productivity, conducted across a national corporation with 80,000 employees. The analysis reveals a statistically significant improvement in output scores (p < 0.0001). What the headline omits is that the average improvement amounts to 0.3 additional tasks completed per week—a difference that disappears entirely against the backdrop of normal day-to-day variation in any individual's workload. The p-value confirmed that the effect was real; it said nothing about whether the effect was worth caring about.

This distinction between statistical significance and practical significance is not a minor technical footnote. It has real consequences for policy decisions, clinical guidelines, educational interventions, and resource allocation.

Effect Sizes: Giving Significance Its Context

Effect sizes are the natural complement to significance testing because they quantify the magnitude of a relationship or difference, independent of sample size. Where a p-value answers the question "is this effect real?", an effect size answers the question "how large is it?"

The most commonly used effect size metrics in behavioral and social science research are Cohen's d for comparing two means, Pearson's r for correlations, and odds ratios or risk ratios in epidemiological contexts. Cohen's widely cited benchmarks—small (d = 0.2), medium (d = 0.5), and large (d = 0.8)—provide a rough interpretive framework, though researchers are increasingly encouraged to calibrate their interpretation against effect sizes typical in their specific field rather than applying universal labels.

Reporting effect sizes alongside p-values is now a formal requirement in many high-impact journals, including those published by the American Psychological Association. If your current manuscripts do not include them, incorporating effect sizes is one of the most immediate ways to strengthen the credibility and interpretability of your work.

Confidence Intervals: Communicating Uncertainty Honestly

Confidence intervals offer a third layer of information that neither p-values nor point estimates alone can provide: a quantified range of uncertainty around your estimate. A 95% confidence interval, properly interpreted, indicates that if you repeated your study many times under identical conditions, approximately 95% of the intervals you calculated would contain the true population parameter.

In practical terms, a confidence interval tells readers how precisely your study has estimated an effect. A narrow interval suggests your estimate is stable and replicable. A wide interval signals that your data support a broad range of possible true values, and that caution is warranted before drawing strong conclusions.

Consider two studies both reporting that a reading intervention improved test scores by 4 points. Study A reports a 95% confidence interval of [3.1, 4.9]. Study B reports a confidence interval of [-1.2, 9.2]. Both studies found the same point estimate, but they tell fundamentally different stories about the reliability of that estimate. Study B's interval crosses zero, meaning the data are consistent with no effect at all. Reporting only the point estimate—or worse, only the p-value—would obscure that crucial distinction.

The Replication Crisis as a Methodological Wake-Up Call

The conversation about p-values has gained urgency in the context of the broader replication crisis, which has called into question the reliability of landmark findings across psychology, nutrition science, cancer biology, and economics. A 2015 large-scale replication effort coordinated by the Center for Open Science found that fewer than half of 100 published psychology studies produced results consistent with the originals when independent teams attempted to reproduce them.

While the causes of the replication crisis are multiple and complex, over-reliance on p-value thresholds is widely recognized as a contributing factor. When researchers face pressure to produce significant results—whether from advisors, journals, or funding bodies—the temptation to engage in practices like selective outcome reporting, flexible stopping rules, or running multiple analyses until p < 0.05 is achieved becomes structurally embedded in the research process. These practices, sometimes grouped under the term "p-hacking," inflate the false positive rate in the published literature and undermine the cumulative knowledge base that subsequent research is built upon.

Toward a More Honest Statistical Language

Shifting away from significance-only reporting is not simply a technical adjustment—it requires a change in how researchers think about and communicate the value of their findings. A result that fails to reach p < 0.05 but demonstrates a consistent, moderate effect size with a reasonably narrow confidence interval may be far more scientifically valuable than a "significant" finding driven by an enormous sample and a negligible effect.

Practical steps researchers can take immediately include: reporting effect sizes as a standard element of results sections; presenting confidence intervals rather than relying solely on significance statements; describing findings in terms of their substantive meaning, not just their statistical status; and avoiding language like "proved" or "confirmed" in favor of more precise characterizations of what the data do and do not support.

Statistical tools are exactly that—tools. Like any instrument, their value depends entirely on whether they are used with an understanding of what they measure and what they cannot. Developing that understanding is not optional for researchers who want their work to contribute meaningfully to their fields. It is a core professional competency, and one that the Research Skill Center is committed to helping every researcher build.

All Articles

Related Articles

7 Checkpoints That Will Strengthen the Reproducibility of Your Research—Starting Today

7 Checkpoints That Will Strengthen the Reproducibility of Your Research—Starting Today

Collecting Less, Understanding More: A Researcher's Guide to Purposeful Study Design

Collecting Less, Understanding More: A Researcher's Guide to Purposeful Study Design

Beyond the First Page of Results: How to Build a Literature Review That Misses Nothing

Beyond the First Page of Results: How to Build a Literature Review That Misses Nothing