Field note · August 22, 2023
We p-hacked ourselves and caught it in review
Fourteen segments, one significant result, and the correction that made it go away.
An experiment came back flat. Primary metric moved 0.2 points, interval spanning zero, clearly nothing.
The team then did what everyone does: sliced it. Device, country, tenure band, plan, acquisition channel, new versus returning, and a couple of combinations. Fourteen tests.
Tablet users on the annual plan showed a 6.1% lift at p = 0.031.
The write-up was drafted as "no overall effect, but a strong effect for annual tablet users — recommend targeted rollout". Everything in that sentence is technically defensible and the conclusion is almost certainly false.
Fourteen tests at α = 0.05 gives you roughly a 51% chance of at least one significant result with nothing real happening. Holm–Bonferroni on those fourteen p-values leaves nothing standing.
What we changed, and it is process rather than statistics:
The primary metric is declared in the experiment doc before launch. One metric. It takes one line.
Segment analyses are labelled exploratory in the template itself, so the label is the default rather than something someone has to remember to add.
The write-up states how many tests were run. Even without a correction, "we looked at fourteen segments" lets a reader discount appropriately. Omitting that number is the part that is actually dishonest.
Nobody was cheating. The team was being thorough, and thoroughness without a pre-registered plan is indistinguishable from fishing. That is the uncomfortable part.