‘Statistically Significant’ Isn’t the Same as ‘Worth Doing’
You are three quarters through a case study. A chart, a percentage, then the line that closes the section: the result was statistically significant. Something in you relaxes. But that phrase answers a narrower question than the one you were asking.
The phrase that ends the argument before it starts
Watch where it gets placed. Rarely beside the methodology. Almost always in the last line, right where a careful reader would start asking harder questions: how big was the lift, over what window, at what cost. It arrives first and functions like a QED. That works because of a word collision. In ordinary English, significant means important, large. In statistics it means something almost unrelated: probably not a fluke. Most readers never make the switch: they see “significant”, read “big”, and stop evaluating.
What “statistically significant” actually claims
The machinery compares your result against a boring alternative: nothing real is happening, and the gap between groups is ordinary sampling wobble. That assumption is the null hypothesis. The NIST Engineering Statistics Handbook describes a statistical test as deciding whether there is enough evidence to reject it, the significance level being the accepted risk of rejecting it wrongly.
The usual output is a p-value: how often a pattern this extreme would appear if nothing were going on. “Significant” means that number came in under a threshold chosen beforehand, usually 0.05. Wikipedia’s article on statistical significance records that Fisher proposed one in twenty as a convenient cutoff in 1925 and later argued levels should suit the circumstances, and that particle physics uses stricter ones. It is a convention, not a law, and all of it is a statement about chance, never about size.
What it does not claim
- That the effect is large. Wikipedia puts it flatly: a statistically significant result may have a weak effect, which is why effect size should be reported too.
- That the effect is worth its cost. The test knows nothing about what the campaign cost or a conversion is worth.
- That the effect will persist. It describes only the window it ran on.
- That a cause has been found. Only that a pattern was unlikely to be random given the sample. What produced it is a separate argument the case study still owes you.
- That the test was well built. Significance is computed on whatever data it is handed. It cannot tell you the test was stopped honestly, the groups were comparable, or the window was not picked afterwards.
A hypothetical case: a tiny lift, a huge sample, a clean significant result
The numbers below are invented to show the mechanism. Not a real reported result, not any agency’s work, not this site’s data.
Picture a brand testing a landing page on a very large paid campaign, 400,000 visitors per variant. The control converts at 2.00 percent, the challenger at 2.12 percent. That 0.12 of a percentage point would vanish into noise on a few thousand visitors. At this scale it does not: the difference runs several standard errors wide and clears a 0.05 threshold easily.
Nothing has gone wrong. The test behaved as designed: the lift is real, not noise. That is all it established. The lift comes to roughly 480 extra conversions, and whether that is a triumph or a rounding error depends on two numbers the case study never gave you: what a conversion is worth and what the campaign cost. A very large sample can carry a very small effect across a line that was never measuring size.
The other direction: a real, meaningful effect that never reaches significance
“Not significant” is not a finding that the effect is zero. It describes what the test could see at that sample size.
The NIST handbook is direct: failing to reject the null may simply mean you do not yet have enough data, and small discrepancies are harder to detect than large ones. Wikipedia’s article on statistical power says it forwards: power is the probability of detecting an effect that genuinely exists, and more data tends to provide more power. A campaign that moved something substantially can still report “not significant” if it ran on a few hundred people. That is inconclusive, not absent. It is the mirror image of when the base is small, the percentage is just noise: there a tiny denominator inflates a percentage, here a tiny sample hides a real effect.
The two questions a case study owes you, not one
Split the claim in half. First, the noise question: given this sample, is the pattern likely to be something other than random variation? Significance testing answers exactly that. Second, the size question, sometimes called practical significance: is the effect big enough, against what it cost, to be worth repeating? No p-value touches it.
A case study that answers the first and stops gives you half of what you came for, while sounding like all of it. This habit sits alongside what a case study leaves out when it names the winning variant and how to read a marketing case study without being misled.
A checklist for the next “statistically significant” claim you read
- Does it state the effect size in real units? Percentage points, conversions, revenue, not a relative percentage with no absolute behind it.
- Does it state the sample size, per variant rather than in total?
- Does it state or imply the threshold used? If the number is never named, you are trusting a cutoff you cannot see.
- Does it keep “this is probably real” and “this is worth doing” as separate sentences?
- Is the effect weighed against a cost or a business target, or only against the null hypothesis?
Read a few with this in hand
Take both questions to a stack of published work. Browse the case study library, each entry credited to the agency that reported it, or submit your own.
FAQ
If a case study says the result was statistically significant, does that mean the campaign worked?
Not on its own. It means the pattern is unlikely to be random variation, given the sample size and threshold. Whether it worked depends on the effect’s size and cost, which significance testing does not measure.
Does “not statistically significant” mean the campaign failed?
No. The test did not detect an effect at the sample size it had, and a small sample can miss an effect that is real and large. Read it as inconclusive and go find the sample size.