From Fragile to Fortress: A Researcher's Practical Guide to Designing Studies That Withstand Scrutiny
Why Your Study Might Not Survive the Next Attempt to Replicate It
Every researcher operates under an implicit assumption: that their findings, if sound, could be reproduced by an independent team working under similar conditions. Yet decades of evidence suggest this assumption is far more fragile than the scientific community once believed. Across psychology, biomedicine, economics, and the social sciences, replication rates in some fields hover uncomfortably below 50 percent. The question for working researchers is no longer whether the reproducibility problem exists—it is whether their own study design is contributing to it.
Understanding the structural causes of replication failure is the first step toward building research that holds up. The second, more demanding step is translating that understanding into concrete design choices before data collection begins. This article addresses both.
The Architecture of a Fragile Study
Replication failures rarely stem from outright fraud. More commonly, they emerge from a cluster of practices that individually seem defensible but collectively erode the reliability of findings.
Underpowered samples remain one of the most pervasive culprits. A study that recruits 40 participants to detect a small-to-medium effect is not merely underpowered—it is structurally optimistic. When effects are real but modest, small samples produce inflated effect size estimates that subsequent studies, drawing from larger or more representative populations, cannot reproduce. A formal a priori power analysis, using conservative effect size estimates derived from prior meta-analyses rather than the most impressive pilot data, is a non-negotiable starting point.
Researcher degrees of freedom—the term coined by psychologists Uri Simonsohn, Leif Nelson, and Joseph Simmons—describe the many branching decision points that occur between raw data and published result. Which outliers to exclude, which covariates to include, whether to analyze subgroups, how to operationalize a construct: each decision, made post hoc and selectively reported, inflates the false-positive rate in ways that are invisible to readers. A 2011 simulation demonstrated that flexible analytic choices alone could push a study's effective Type I error rate from 5 percent to over 60 percent.
Selective outcome reporting compounds this problem. When researchers register five outcomes and report only the two that reached significance, the published result is not a finding—it is a selection artifact. The 2013 AllTrials campaign in the clinical sciences highlighted how widespread this practice was in pharmaceutical research, with consequences that extended well beyond academic reputation.
Case Studies in Replication: What the Evidence Shows
The 2015 Reproducibility Project, coordinated by the Center for Open Science, attempted to replicate 100 studies published in three prominent psychology journals. Only 36 percent of replications produced statistically significant results in the same direction as the original. Crucially, effect sizes in the replications were, on average, half the magnitude of the originals—consistent with the hypothesis that small samples and flexible analyses had inflated initial estimates.
Contrast this with the Registered Replication Reports, a format in which methodological details are pre-approved before data collection and multiple independent labs collect data simultaneously. Studies conducted under this framework have demonstrated markedly higher replication rates, suggesting that the problem is not inherent to the phenomena being studied but to the practices surrounding how those phenomena are investigated.
In biomedical research, a 2012 report by scientists at Amgen found that only 11 of 53 landmark cancer biology studies could be reproduced internally—a finding that prompted significant reflection on preclinical research standards and eventually contributed to the development of the Reproducibility Project: Cancer Biology.
The Pre-Submission Checklist: Designing for Durability
The following framework is intended to be applied before data collection begins, not as an afterthought during manuscript preparation.
1. Preregister your study. Platforms such as the Open Science Framework (OSF), AsPredicted, and ClinicalTrials.gov allow researchers to publicly record their hypotheses, sample size targets, primary outcomes, and analytic plans before a single observation is made. Preregistration does not eliminate exploratory analysis—it distinguishes confirmatory from exploratory findings, which is epistemically essential.
2. Conduct a rigorous power analysis. Use G*Power or equivalent software. Base effect size estimates on published meta-analyses or, where none exist, on theoretically conservative assumptions. Aim for 80 percent power at minimum; 90 percent is preferable for primary outcomes.
3. Specify your primary outcome in advance. If your study has one central question, identify one primary outcome measure. Secondary and exploratory outcomes should be clearly labeled as such in both the preregistration and the manuscript.
4. Document your analytic pipeline before unblinding. Write out your data cleaning rules, exclusion criteria, and analysis scripts before examining outcome data. Tools such as R Markdown and Jupyter Notebooks allow researchers to create reproducible, timestamped analytic records that serve as transparent documentation.
5. Share your materials, data, and code. The Open Materials and Open Data badges offered by many journals signal to readers—and future replicators—that verification is possible. Repositories such as OSF, Zenodo, and the Harvard Dataverse provide stable, citable homes for supplementary materials.
6. Report effect sizes and confidence intervals, not just p-values. A p-value below .05 tells a reader that a result is unlikely under the null hypothesis. An effect size with a confidence interval tells them how large the effect probably is and how precisely it was estimated. The latter is far more useful to anyone attempting to build on your work.
7. Conduct a sensitivity analysis. Test whether your primary conclusions hold under alternative analytic assumptions—different outlier thresholds, alternative covariate sets, or different operationalizations of your key variables. Findings that evaporate when a single decision changes were never as robust as they appeared.
Transparency as a Methodological Value
Beyond checklists, designing reproducible research requires a shift in orientation. The goal of a study is not to produce a publishable p-value—it is to generate an accurate estimate of a real-world phenomenon. This distinction matters because it changes the decisions researchers make at every stage of the process.
Graduate training programs in the United States have begun incorporating open science practices into their curricula, and funding agencies including the National Institutes of Health and the National Science Foundation have strengthened rigor and reproducibility requirements in grant applications. These are institutional signals worth heeding.
The researchers who will shape the next generation of scientific knowledge are the ones who treat methodological rigor not as a bureaucratic burden but as the foundation upon which credible knowledge is built. Building that foundation is not glamorous work. But it is the only work that lasts.