The p-value of 0.Which means you ran the experiment, crunched the numbers, and there it was — just under the magic line. Publishable. 04 used to feel like a victory. Also, significant. Done Practical, not theoretical..
But 2016 changed how a lot of us look at that number.
That year, Nature didn't just publish research. A series of articles, commentaries, and editorials that forced the scientific community to stare at the p-value and ask: what does this actually mean? It published a reckoning. And more uncomfortably: what have we been getting wrong?
If you've ever stared at a p-value of 0.04 and felt a little uneasy — or if you've reviewed a paper where that number did all the heavy lifting — this is the backstory you need.
What Happened in 2016
The short version: the American Statistical Association (ASA) released a statement on p-values. Nature covered it heavily. And the conversation shifted from "is it significant?" to "what does significance even mean?
The ASA statement — published in The American Statistician but amplified across Nature's news and comment sections — was unprecedented. For the first time, the largest professional organization of statisticians in the world said, publicly and plainly: **stop treating p < 0.05 as a binary truth filter.
Easier said than done, but still worth knowing.
Nature ran multiple pieces that year. A news story titled "Statisticians issue warning over misuse of P values." A commentary from David Colquhoun arguing that p = 0.04 means you're still 26% likely to be wrong (under reasonable priors). An editorial urging journals to rethink statistical standards.
It wasn't a single article. It was a campaign.
And the p-value of 0.04 became a kind of mascot for the problem — close enough to 0.05 to feel "significant," far enough from zero to be deeply unstable.
Why 0.04 Is the Perfect Troublemaker
Let's be honest: nobody loses sleep over p = 0.06 gets rejected without a second glance. Which means 04? And p = 0.001. But 0.That's the danger zone.
It's the "just significant" result. Consider this: the one that gets highlighted in the abstract. The one that makes it into the press release. The one that feels like evidence.
But here's what 2016 made clear: a p-value of 0.04, in a typical underpowered study, can easily correspond to a false positive rate of 20–30% or higher.
That's not a typo. Here's the thing — it's not a controversial take. It's basic probability The details matter here..
The p-value tells you: if the null hypothesis is true, how likely is this data (or more extreme)? It does not tell you: how likely is the null hypothesis true given this data? That's a different question — one that requires prior probabilities, power, and a Bayesian framework.
Honestly, this part trips people up more than it should Easy to understand, harder to ignore..
Most researchers don't think in those terms. In 2016, Nature made sure more of us started to.
The Colquhoun Argument
David Colquhoun, a pharmacologist at UCL, wrote a blistering piece for Nature in 2016 (and a longer paper in Royal Society Open Science) showing that if you test a hypothesis with a prior probability of 10% — which is generous for many exploratory fields — a p-value of 0.04 gives you a false positive risk of about 26%.
Real talk — this step gets skipped all the time.
Read that again. One in four.
And that's assuming no p-hacking, no selective reporting, no researcher degrees of freedom. Worth adding: in the real world? The false positive risk is almost certainly higher Easy to understand, harder to ignore. Nothing fancy..
This wasn't new math. But Nature gave it a platform. And suddenly, that comfortable 0.04 started looking a lot shakier Simple, but easy to overlook..
The ASA's Six Principles
The ASA statement didn't just criticize. It offered guidance. Six principles, widely cited, widely ignored, and absolutely worth revisiting:
-
P-values can indicate how incompatible the data are with a specified statistical model. That's it. Not "the probability the hypothesis is true." Not "the probability the result is real." Just incompatibility with a model.
-
P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone. This is the most common misinterpretation. Full stop Most people skip this — try not to..
-
Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold. The 0.05 cliff is arbitrary. Always has been.
-
Proper inference requires full reporting and transparency. No cherry-picking. No hiding the three analyses that didn't work And that's really what it comes down to..
-
A p-value, or statistical significance, does not measure the size of an effect or the importance of a result. A tiny effect with a huge sample can yield p = 0.0001. A massive effect with a tiny sample can yield p = 0.2. Which one matters?
-
By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis. It's one piece. Not the verdict.
Nature published explainers on each of these. They ran opinion pieces from statisticians, psychologists, geneticists, and journal editors. The message was consistent: the p-value is a tool, not a truth machine.
What Nature Actually Published That Year
It wasn't just news coverage. Nature published primary research and commentary that shaped the conversation:
- The ASA statement (summarized and contextualized in Nature News, March 2016)
- Editorial: "Statisticians' warning on p-values" — calling for cultural change, not just technical fixes
- Comment: "The fickle P value generates irreproducible results" — by Regina Nuzzo, one of the clearest explainers ever written on this topic
- Correspondence from researchers across fields — some defensive, some relieved, some demanding more
- A special collection on statistical challenges in irreproducible research
And then there was the Nuzzo piece. Because of that, if you read one thing from 2016 on this, make it that. Even so, she walked through the history of the p-value, the Fisher-Neyman-Pearson wars, the rise of NHST (null hypothesis significance testing), and the quiet disaster of treating p < 0. 05 as a publishability gate.
The official docs gloss over this. That's a mistake.
She also introduced a lot of readers to the concept of false positive risk — the probability that a "significant" result is actually a fluke. Not the p-value. The actual risk
of being wrong. That distinction—between a p-value and the false positive risk—is where the real confusion lies, and where the reform movement gained its sharpest edge Surprisingly effective..
Nuzzo’s article didn’t just explain the problem; it offered a way forward. She urged researchers to stop treating p-values as the be-all and end-all and instead see them as one piece of a larger puzzle. Even so, she argued that p-values, when properly contextualized, can still be useful. But they must be reported alongside other information: effect sizes, confidence intervals, sample sizes, and study design. Her message resonated because it didn’t demand the abandonment of statistics—it demanded better statistics Less friction, more output..
This is the bit that actually matters in practice.
The response was swift. Journals began revising their guidelines. Some, like Nature and The Lancet, started explicitly discouraging the sole reliance on p-values. In real terms, others, like Psychological Science, introduced reforms such as requiring the reporting of effect sizes and confidence intervals alongside p-values. Funding agencies, too, began to ask for more nuanced statistical reporting in grant proposals. The tide was turning, but slowly Simple as that..
Yet, the cultural shift proved harder than the technical one. Still, decades of training in NHST—where a p-value below 0. Even so, habits of mind are stubborn. In some fields, particularly medicine and psychology, the pressure to publish in high-impact journals still incentivized the pursuit of “significant” results. 05 was a golden ticket—didn’t vanish overnight. Researchers, reviewers, and editors alike often defaulted to the old ways, even as they paid lip service to the new reforms.
The reproducibility crisis, which had been simmering for years, exploded in 2015 with the Open Science Collaboration’s replication project. The result? Now, over 200 labs attempted to replicate 100 psychological studies. Because of that, only 39% of the effects held up. The failure rate was staggering. The replication crisis became a rallying cry for change, and the ASA’s 2016 statement, amplified by Nature, became a key document in that movement Easy to understand, harder to ignore..
But even as the conversation grew louder, resistance persisted. Others, particularly in fields like genomics and physics, defended the continued use of p-values, especially when adjusted for multiple comparisons. Consider this: they pointed out that the problem wasn’t the p-value itself, but its misuse. Some statisticians argued that p-values were being unfairly vilified. The debate was far from settled.
What emerged, however, was a consensus: the p-value, in its traditional form, was no longer sufficient. Practically speaking, it needed to be part of a broader, more transparent approach to statistical analysis. Practically speaking, it called for preregistration of studies, open data, and replication efforts. On the flip side, this approach emphasized not just the result, but the process—how the data were collected, analyzed, and interpreted. It demanded that science stop chasing statistical significance and start seeking meaningful understanding.
The Nature explainers and ASA statement were not the end of the story, but a turning point. Here's the thing — they marked the beginning of a long-overdue reckoning with how we do science. The p-value, once a symbol of scientific authority, was now being reimagined as a tool—one that, when used responsibly, could help guide discovery without dictating it That's the part that actually makes a difference..
In the end, the message was clear: science is not about finding a magic number. Consider this: it’s about asking better questions, collecting better data, and thinking more critically about what the numbers really mean. The p-value, when properly understood, can be part of that process—but only if we stop treating it as the final word.