Why a Test Stopped Early Overstates the Winning Margin

Why a Test Stopped Early Overstates the Winning Margin

A case study tells you a variant beat the control by a wide margin. It does not tell you when the test was stopped, or who decided to stop it. That missing detail is the difference between a finding and a snapshot of the luckiest week the test had.

The habit that breaks the test

A test goes live Monday. Someone opens the dashboard Wednesday, sees the new version ahead, checks Thursday, sees it still ahead, and calls it. Nothing was faked. The split was clean and the tracking worked. It ended at the most flattering moment available, and the moment was chosen by looking.

That is the whole mechanism: not a bad test, a good test ended opportunistically. Statisticians have several names for it, peeking, optional stopping, repeated significance testing, so this is not a novel accusation. Evan Miller’s note sets out the arithmetic. A dashboard’s significance figure assumes the sample size was fixed in advance, and if you instead run until it looks significant, that figure becomes meaningless. His worked case, a 50 percent baseline checked after every observation and called off at 150 observations, produces a wrong significant result 26.1 percent of the time when the change does nothing.

Why an early lead is not a stable lead

Two things are happening here. Keep them apart. The first is the size of the number. Early on you have few observations, so a handful of conversions in one bucket swings the percentage hard. An extreme early reading is more likely to be extreme partly by luck, and luck does not persist. As observations accumulate the estimate tightens around the real difference. That is regression to the mean applied to a running estimate. The underlying effect is not shrinking. Nothing about the variant got worse. The measurement got better, and the flattering part of the early number was never real. Same reason when the base is small, the percentage is noise.

The second is whether the result counts as significant, and it differs from p-hacking in general. Ordinary p-hacking slices the data many ways until one slice looks good. Optional stopping checks one metric many times. Every check is another chance for noise to cross the threshold, and stopping on one of them locks in that moment. Miller quantifies it: peek ten times and what reads as 1 percent significance is really 5 percent.

A hypothetical test, day by day

These numbers are invented to illustrate a shape. They are not a reported result from any real test, in this library or anywhere else.

Checkpoint Sessions per variant Reported lift
Day 3 310 38 percent
Day 7 840 19 percent
Day 14 1,760 11 percent
Day 21 2,640 7 percent

Stopped on day 3, this hypothetical test publishes a 38 percent lift. Run to day 21, the same test, same variant, same audience, publishes 7 percent. Both sat on a dashboard at some point. Only one is a result. A reader handed the day 3 version cannot tell it apart from a test that genuinely ran three weeks, because case studies print the percentage and not the timeline.

What a case study almost never tells you about its own test

Three disclosures would settle the question:

  • The planned duration or sample size, fixed before the test went live. Not “a few weeks”, a number that existed before any data did.
  • The duration or sample size the test actually reached. If the plan said 4,000 sessions per arm and the test reported at 900, that gap is the story.
  • Whether the stopping point was decided before or after seeing the result. The load-bearing one, and almost nobody volunteers it.

What you get instead is a single final percentage: no dates, no sample size, no pre-set stopping rule, no statement of who called it. The absence is so standard it barely reads as one.

The tell in the writing, not the number

You will rarely get the daily data, so read the prose. Speed presented as a virtue is the clearest signal. Phrases like “within days”, “immediately outperformed”, or “results were so strong we rolled it out early” describe an opportunistic stopping point in the language of a triumph. The writer is not hiding it, they are proud of it.

The opposite tell is dull and specific. A test window with real dates. A stated sample size target. A significance threshold set before launch, with the test run to it. A clause saying the test was planned for four weeks and reported at four weeks does more work than any percentage on the page.

What a trustworthy version of this claim looks like

  • The stopping rule came first. Duration or sample size set before the test started, not selected once early results were visible.
  • The reported margin is the one measured at the pre-set stopping point. Not the widest gap the test showed at any moment.
  • The test collected enough to let a swing settle. Enough observations, or enough calendar time across normal weekly variation, that an early bounce would have averaged out.

All three fit in two sentences. A case study making none of them is asking you to trust a number with its most important property left out.

Take the question to the library

Fluency comes from reading published work with the question in hand. Ten case studies are published here, each credited to the agency that produced it, and every number belongs to that agency. Browse the case study library and ask each test claim when it stopped and who decided. If your agency has work you would like read this closely, submit your own. For the wider framework, see how to read a marketing case study.

FAQ

Is stopping a test early always dishonest?

No. Most people who stop early are not lying, they are checking a dashboard and reacting to what they see, as anyone would. The problem is the statistics, not the intent: checking repeatedly and stopping on a good reading inflates the result whether or not anybody meant it to.

Does a longer test always mean a more trustworthy result?

Not automatically. Duration helps, since more observations mean a tighter estimate, but it does not fix a finish line chosen after looking at the data. A test watched daily for six weeks and stopped on a good Tuesday carries the same flaw as one stopped in week one. What matters is whether the stopping rule was set in advance, not how many days elapsed.

Sources

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *