The p-value of 0.Think about it: significant. And publishable. 04 used to feel like a victory. You ran the experiment, crunched the numbers, and there it was — just under the magic line. Done.
But 2016 changed how a lot of us look at that number.
That year, Nature didn't just publish research. But it published a reckoning. A series of articles, commentaries, and editorials that forced the scientific community to stare at the p-value and ask: what does this actually mean? And more uncomfortably: what have we been getting wrong?
Some disagree here. Fair enough That's the part that actually makes a difference..
If you've ever stared at a p-value of 0.04 and felt a little uneasy — or if you've reviewed a paper where that number did all the heavy lifting — this is the backstory you need Small thing, real impact..
What Happened in 2016
The short version: the American Statistical Association (ASA) released a statement on p-values. Nature covered it heavily. And the conversation shifted from "is it significant?" to "what does significance even mean?
The ASA statement — published in The American Statistician but amplified across Nature's news and comment sections — was unprecedented. So for the first time, the largest professional organization of statisticians in the world said, publicly and plainly: **stop treating p < 0. 05 as a binary truth filter The details matter here..
Nature ran multiple pieces that year. A news story titled "Statisticians issue warning over misuse of P values." A commentary from David Colquhoun arguing that p = 0.04 means you're still 26% likely to be wrong (under reasonable priors). An editorial urging journals to rethink statistical standards.
It wasn't a single article. It was a campaign.
And the p-value of 0.04 became a kind of mascot for the problem — close enough to 0.05 to feel "significant," far enough from zero to be deeply unstable.
Why 0.04 Is the Perfect Troublemaker
Let's be honest: nobody loses sleep over p = 0.That's why 06 gets rejected without a second glance. And p = 0.04? 001. But 0.That's the danger zone.
It's the "just significant" result. Day to day, the one that makes it into the press release. The one that gets highlighted in the abstract. The one that feels like evidence.
But here's what 2016 made clear: a p-value of 0.04, in a typical underpowered study, can easily correspond to a false positive rate of 20–30% or higher.
That's not a typo. In real terms, it's not a controversial take. It's basic probability.
The p-value tells you: *if the null hypothesis is true, how likely is this data (or more extreme)?Even so, * It does not tell you: *how likely is the null hypothesis true given this data? * That's a different question — one that requires prior probabilities, power, and a Bayesian framework Practical, not theoretical..
Most researchers don't think in those terms. In 2016, Nature made sure more of us started to.
The Colquhoun Argument
David Colquhoun, a pharmacologist at UCL, wrote a blistering piece for Nature in 2016 (and a longer paper in Royal Society Open Science) showing that if you test a hypothesis with a prior probability of 10% — which is generous for many exploratory fields — a p-value of 0.04 gives you a false positive risk of about 26%.
Read that again. One in four That's the part that actually makes a difference..
And that's assuming no p-hacking, no selective reporting, no researcher degrees of freedom. Practically speaking, in the real world? The false positive risk is almost certainly higher.
This wasn't new math. But Nature gave it a platform. And suddenly, that comfortable 0.04 started looking a lot shakier.
The ASA's Six Principles
The ASA statement didn't just criticize. It offered guidance. Six principles, widely cited, widely ignored, and absolutely worth revisiting:
-
P-values can indicate how incompatible the data are with a specified statistical model. That's it. Not "the probability the hypothesis is true." Not "the probability the result is real." Just incompatibility with a model.
-
P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone. This is the most common misinterpretation. Full stop Small thing, real impact. Worth knowing..
-
Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold. The 0.05 cliff is arbitrary. Always has been Worth keeping that in mind..
-
Proper inference requires full reporting and transparency. No cherry-picking. No hiding the three analyses that didn't work.
-
A p-value, or statistical significance, does not measure the size of an effect or the importance of a result. A tiny effect with a huge sample can yield p = 0.0001. A massive effect with a tiny sample can yield p = 0.2. Which one matters?
-
By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis. It's one piece. Not the verdict.
Nature published explainers on each of these. They ran opinion pieces from statisticians, psychologists, geneticists, and journal editors. The message was consistent: the p-value is a tool, not a truth machine.
What Nature Actually Published That Year
It wasn't just news coverage. Nature published primary research and commentary that shaped the conversation:
- The ASA statement (summarized and contextualized in Nature News, March 2016)
- Editorial: "Statisticians' warning on p-values" — calling for cultural change, not just technical fixes
- Comment: "The fickle P value generates irreproducible results" — by Regina Nuzzo, one of the clearest explainers ever written on this topic
- Correspondence from researchers across fields — some defensive, some relieved, some demanding more
- A special collection on statistical challenges in irreproducible research
And then there was the Nuzzo piece. If you read one thing from 2016 on this, make it that. That said, she walked through the history of the p-value, the Fisher-Neyman-Pearson wars, the rise of NHST (null hypothesis significance testing), and the quiet disaster of treating p < 0. 05 as a publishability gate.
She also introduced a lot of readers to the concept of false positive risk — the probability that a "significant" result is actually a fluke. Not the p-value. The actual risk
of being wrong. That distinction—between a p-value and the false positive risk—is where the real confusion lies, and where the reform movement gained its sharpest edge.
Nuzzo’s article didn’t just explain the problem; it offered a way forward. But they must be reported alongside other information: effect sizes, confidence intervals, sample sizes, and study design. Still, she argued that p-values, when properly contextualized, can still be useful. She urged researchers to stop treating p-values as the be-all and end-all and instead see them as one piece of a larger puzzle. Her message resonated because it didn’t demand the abandonment of statistics—it demanded better statistics.
The response was swift. Plus, journals began revising their guidelines. Some, like Nature and The Lancet, started explicitly discouraging the sole reliance on p-values. Others, like Psychological Science, introduced reforms such as requiring the reporting of effect sizes and confidence intervals alongside p-values. Funding agencies, too, began to ask for more nuanced statistical reporting in grant proposals. The tide was turning, but slowly.
Yet, the cultural shift proved harder than the technical one. Practically speaking, in some fields, particularly medicine and psychology, the pressure to publish in high-impact journals still incentivized the pursuit of “significant” results. Habits of mind are stubborn. Practically speaking, 05 was a golden ticket—didn’t vanish overnight. Decades of training in NHST—where a p-value below 0.Researchers, reviewers, and editors alike often defaulted to the old ways, even as they paid lip service to the new reforms.
The reproducibility crisis, which had been simmering for years, exploded in 2015 with the Open Science Collaboration’s replication project. The result? Consider this: over 200 labs attempted to replicate 100 psychological studies. Only 39% of the effects held up. The failure rate was staggering. The replication crisis became a rallying cry for change, and the ASA’s 2016 statement, amplified by Nature, became a key document in that movement Worth keeping that in mind..
But even as the conversation grew louder, resistance persisted. Some statisticians argued that p-values were being unfairly vilified. They pointed out that the problem wasn’t the p-value itself, but its misuse. Others, particularly in fields like genomics and physics, defended the continued use of p-values, especially when adjusted for multiple comparisons. The debate was far from settled That's the part that actually makes a difference..
It sounds simple, but the gap is usually here.
What emerged, however, was a consensus: the p-value, in its traditional form, was no longer sufficient. It called for preregistration of studies, open data, and replication efforts. This approach emphasized not just the result, but the process—how the data were collected, analyzed, and interpreted. It needed to be part of a broader, more transparent approach to statistical analysis. It demanded that science stop chasing statistical significance and start seeking meaningful understanding And that's really what it comes down to..
Some disagree here. Fair enough The details matter here..
The Nature explainers and ASA statement were not the end of the story, but a turning point. They marked the beginning of a long-overdue reckoning with how we do science. The p-value, once a symbol of scientific authority, was now being reimagined as a tool—one that, when used responsibly, could help guide discovery without dictating it.
In the end, the message was clear: science is not about finding a magic number. Think about it: it’s about asking better questions, collecting better data, and thinking more critically about what the numbers really mean. The p-value, when properly understood, can be part of that process—but only if we stop treating it as the final word.