Why the Best Metric on a Dashboard Is Rarely the Real Story
You are reading a case study and one number is doing all the work. Look at the dashboard behind it first. If it tracked twenty metrics, at least one was fairly likely to look like a win whatever the campaign did.
The mechanism: why tracking more numbers finds more wins
The name for this is the multiple comparisons problem. In statistics it covers analyses running many tests at once: compare two groups on enough attributes and it becomes increasingly likely they differ on at least one through sampling error alone. Searching data afterwards for any notable pattern, with no hypothesis fixed in advance, has its own name, data dredging, or p-hacking when applied to significance tests.
Dashboards run no significance tests, so the vocabulary does not transfer cleanly. The arithmetic does. Every metric moves week to week for reasons unrelated to the campaign, seasonality or a post a stranger shared. Watch twenty of them and some will move a long way in the favourable direction. That is not an accusation. The trap needs nobody to lie, or even to notice.
How this shows up when a case study gets written
The trap closes when someone opens the reporting after the campaign and decides what the story is. Two kinds of metric are available then, and the finished document rarely distinguishes them. The first was named as the target before anything ran, in a brief or scope of work that predates the result. The second is one nobody set out to move that moved anyway.
A working dashboard carries reach, impressions, engagement rate, click through rate, conversion rate, cost per result, follower growth, video completion, saves, shares and more. Each is another chance for the second kind to appear. The format does the rest: one headline number plus a short narrative rewards the best available figure, not the predicted one.
A hypothetical dashboard, worked through
Here is a hypothetical campaign with hypothetical numbers, invented for this post and drawn from no published case study, agency or brand. Picture a regional furniture retailer running a four week push. The stated goal in the hypothetical brief was cost per result, and twenty metrics were tracked.
| Hypothetical metric | Hypothetical four week change |
|---|---|
| Cost per result (stated goal) | 2% worse |
| Reach | up 4% |
| Impressions | up 6% |
| Engagement rate | down 1% |
| Click through rate | up 3% |
| Conversion rate | flat |
| Follower growth | up 5% |
| Video completion | down 2% |
| Saves | up 240% |
| Shares | up 7% |
| Comments | down 3% |
| Profile visits | up 8% |
| Link clicks | up 2% |
| Cost per click | 4% worse |
| Cost per 1,000 impressions | 1% better |
| Average watch time | flat |
| Story replies | down 9% |
| Direct messages | up 5% |
| Mentions | down 6% |
| Branded search | flat |
Nineteen of these hypothetical numbers are noise. One is a headline, and it is not the metric the brief named: cost per result, the thing the campaign was built to move, went slightly the wrong way. Reported honestly, the same campaign reads differently: cost per result was the target and worsened by 2 percent over a fixed four week window, while saves rose sharply on a low base, with no claim about why. Every number above is hypothetical.
Four questions that separate a predicted result from a found one
- Was this metric named as the goal before the campaign started? A write-up naming the objective agreed in the brief makes a checkable claim. One that opens with the number and never says what was aimed at does not.
- Are the other tracked metrics reported too? If only one appears, you cannot tell whether it was one of three or one of thirty.
- Would the result hold over a different window? Ask whether the period was fixed in advance or chosen once the data existed. A result that appears in one four week slice and vanishes across the quarter is a window, not an outcome.
- Is the move plausible for the spend and audience? A very large percentage on a small base is arithmetic, not achievement. If it outruns what was actually done, ask what the base was.
Why adding more tracked metrics makes the problem worse, not better
Each additional metric is one more chance for ordinary variation to throw up a good-looking number. Three metrics gives three chances. Twenty gives twenty. With identical underlying performance, the twenty-metric dashboard is more likely to hold a standout figure, because it holds more figures.
That runs against the instinct that a longer dashboard is a more rigorous one. For this purpose it is the opposite, unless the target was fixed in advance. Twenty metrics with no stated target is a menu. See also how to read a marketing case study without being misled and why the bad quarter flatters the next.
What a credible case study discloses that this trap hides
- The target metric, named as a target. Not “we improved engagement rate” but “engagement rate was the agreed objective, set before launch”.
- The full tracked set, not the winner. Everything measured, listed, so the highlighted number can be read against the rest.
- The measurement window and when it was set. Explicit start and end dates, and whether they were fixed before the campaign ran.
- At least one metric that did not move, or moved the wrong way, where that is true. A write-up where every reported number improved is either an unusually clean campaign or a filtered one.
None of that needs a longer document, only a target committed to before the answer is known. Selection operates one level up too: why every case study library is a selected sample.
Try it on a real write-up
Our ten published case studies are each credited to the agency that produced them, with every number stated as theirs. Browse the case study library, or submit your own.
FAQ
Does this mean every case study with one standout metric is misleading?
No. A standout metric is real when it was the predefined target, or when the full dashboard is disclosed. The trap is narrower: picking the best-looking number after the results are in, then presenting it as though it had been predicted.
What is the difference between data dredging and this dashboard problem?
Data dredging is the statistical term for searching a dataset after the fact for any notable pattern, with no hypothesis fixed beforehand, then reporting it as though it had been expected. Picking the winning metric off a dashboard is an everyday version: the hypothesis came from the data used to confirm it.
Sources
- Multiple comparisons problem, Wikipedia, fetched 30 August 2026.
- Data dredging, Wikipedia, fetched 30 August 2026.