Before You Build on It, Break It: A Researcher's Guide to Auditing Published Studies for Reproducibility
Published does not mean reliable. Peer-reviewed does not mean perfect. These are uncomfortable truths that every serious researcher eventually confronts, often after they have already woven a questionable study into the scaffolding of their own work. The more efficient approach is to develop a disciplined habit of evaluating the reproducibility of any study before it earns a place in your literature review, theoretical framework, or citation list.
This is not about cynicism toward the scientific enterprise. It is about rigor. The same commitment to methodological soundness that you apply to your own research design must extend to the work you choose to build upon. What follows is a structured approach to doing exactly that.
Why Reproducibility Should Be Your First Filter
The reproducibility crisis in scientific research has been well documented across fields ranging from social psychology to biomedical science. A landmark 2015 project coordinated by the Center for Open Science attempted to replicate 100 published psychology studies and found that fewer than half produced results consistent with the original findings. While the magnitude of the problem varies by discipline, no field is entirely immune.
For researchers working in the United States, where grant funding, publication pressure, and academic advancement often create incentives that are not perfectly aligned with methodological caution, this matters practically. When you cite a fragile study, you risk building your argument on a foundation that may shift or collapse. A reproducibility audit—a structured critical review of a study's methods, reporting practices, and transparency—helps you assess that risk before it becomes your problem.
Red Flag One: Sample Size Without Justification
One of the most telling signs of a methodologically underpowered study is a small sample size accompanied by no power analysis or statistical justification. In well-designed research, the sample size is determined before data collection begins, based on an expected effect size and a target level of statistical power—typically 0.80 or higher in most behavioral and social science fields.
When you encounter a study that reports, say, a statistically significant result based on 18 participants with no discussion of how that number was determined, treat it with caution. Small, underpowered samples are more susceptible to chance findings, and statistically significant results from such studies are less likely to replicate.
Look specifically for a methods section that describes a priori power calculations. Their absence is worth noting, though not automatically disqualifying—context matters. Their presence, however, is a meaningful marker of planning and transparency.
Red Flag Two: Vague or Selective Statistical Reporting
A study that reports only whether a finding was statistically significant—without providing the actual test statistics, degrees of freedom, confidence intervals, or effect sizes—is withholding information you need to evaluate its credibility.
Effect sizes, in particular, are essential. A p-value tells you only whether an observed difference is likely to be due to chance; it says nothing about whether that difference is meaningful in a practical or theoretical sense. A study with a very large sample can produce a statistically significant result for an effect so small it carries no real-world relevance. Conversely, a study that fails to report effect sizes makes it nearly impossible for other researchers to conduct meaningful meta-analyses or power their own replications.
When auditing a study, check whether the authors report standardized effect sizes such as Cohen's d, eta-squared, or odds ratios, depending on the design. Check also whether confidence intervals are provided. Selective reporting—where only the significant outcomes are highlighted while non-significant findings are buried or omitted—is another pattern worth watching for, and one that the pre-registration movement has been specifically designed to counteract.
Red Flag Three: Absence of Pre-Registration or Data Availability
Pre-registration, the practice of publicly documenting a study's hypotheses, design, and analysis plan before data collection begins, has become an increasingly important standard in credible empirical research. Platforms such as the Open Science Framework (OSF) and AsPredicted allow researchers to create timestamped records of their intentions, making it easier to distinguish confirmatory from exploratory findings.
A study published in the past several years that makes no mention of pre-registration is not automatically suspect, but it does warrant additional scrutiny. Without pre-registration, there is no reliable way to know whether the hypotheses presented in the paper were formulated before or after the data were examined—a practice sometimes called HARKing, or Hypothesizing After Results are Known.
Similarly, assess whether the authors have made their data and analysis code publicly available. Open data policies are increasingly required by major journals and funding agencies, including the National Institutes of Health and the National Science Foundation. A study whose data cannot be independently examined is considerably harder to verify.
Red Flag Four: Overgeneralized Conclusions
Pay close attention to the distance between what a study actually measured and what its authors claim in the discussion or abstract. A study conducted on a convenience sample of undergraduate students at a single university in the United States should not be the basis for sweeping claims about human behavior broadly defined. When authors make conclusions that extend well beyond the scope of their sample, design, or measures, that gap is itself a warning sign.
This is particularly common in fields where media coverage creates pressure to frame findings in accessible, headline-friendly terms. The abstract may use language that the methods section does not support. Reading both carefully—and comparing them—is a reliable diagnostic practice.
Building a Practical Checklist
To make this process systematic, consider evaluating every key study against the following questions before incorporating it into your work:
- Is the sample size justified with a power analysis or equivalent reasoning?
- Are effect sizes and confidence intervals reported alongside p-values?
- Were the hypotheses pre-registered before data collection?
- Are the raw data or analysis scripts publicly accessible?
- Do the conclusions match the scope and limitations of the design?
- Has the study been independently replicated, or cited in ways that suggest replication attempts?
- Are conflicts of interest, funding sources, and potential biases disclosed?
No single red flag should necessarily disqualify a study from your review. Research is rarely perfect, and older studies predate many of today's transparency norms. The goal is informed judgment, not blanket dismissal. But when multiple warning signs appear together, the appropriate response is skepticism—and a search for corroborating evidence from more methodologically robust sources.
Reproducibility Literacy as a Core Research Skill
Learning to audit published research for reproducibility is not a peripheral skill. It is central to the practice of rigorous scholarship. Every time you choose which studies to trust, you are making a methodological decision that shapes the credibility of your own work. Developing a structured, consistent approach to that decision-making process is one of the most consequential investments you can make as a researcher.