Field note · February 15, 2024

Course note — Describing a column without lying

The mean is the default summary and it is frequently the wrong one. Knowing when — and what to reach for instead — is most of applied statistics in this job.

1 min read ·Statistics ·practice

Revised Describing a column without lying this week. The change was small and the reason was not.

The mean is the default summary and it is frequently the wrong one. Knowing when — and what to reach for instead — is most of applied statistics in this job.

The first draft explained the mechanism and stopped. What was missing was the failure: the thing that happens when you get it wrong, described concretely enough that someone recognises it later. A lesson that only teaches the correct version leaves you unable to spot the incorrect one, which is the situation you will actually be in.

It sits in the Statistics you will actually use module, and the exercise runs against ride-hail-trips. That pairing is deliberate: the dataset was built with the trap the lesson describes already in it, so the exercise fails in the instructive way rather than the confusing one.

Ordering matters here more than in most courses. It follows The first hour with an unfamiliar dataset and leads into Uncertainty, sampling, and how much to trust a number, and reading it out of sequence mostly works but costs you the setup.

Free means free, and it also means we can rewrite it whenever it is wrong. No edition, no errata PDF, nothing to repurchase. The whole course is 36 lessons and the fixes land the day we find them.