Field note · May 6, 2023
Reader question — Describing a column without lying
The mean is the default summary and it is frequently the wrong one. Knowing when — and what to reach for instead — is most of applied statistics in this job.
Reader question on Describing a column without lying, and the answer belongs somewhere more findable than an email.
The mean is the default summary and it is frequently the wrong one. Knowing when — and what to reach for instead — is most of applied statistics in this job.
The question was, roughly, "when does this stop applying?" — which is the right question and the one lessons routinely fail to answer. Every technique has a range of validity, and stating it is what separates a lesson from a recipe.
It sits in the Statistics you will actually use module, and the exercise runs against ride-hail-trips. That pairing is deliberate: the dataset was built with the trap the lesson describes already in it, so the exercise fails in the instructive way rather than the confusing one.
Ordering matters here more than in most courses. It follows The first hour with an unfamiliar dataset and leads into Uncertainty, sampling, and how much to trust a number, and reading it out of sequence mostly works but costs you the setup.
Free means free, and it also means we can rewrite it whenever it is wrong. No edition, no errata PDF, nothing to repurchase. The whole course is 36 lessons and the fixes land the day we find them.