Field note · June 4, 2025

Cutting a lesson down — Describing a column without lying

The mean is the default summary and it is frequently the wrong one. Knowing when — and what to reach for instead — is most of applied statistics in this job.

1 min read ·Statistics ·practice

A note from writing Describing a column without lying, which took three passes to get to something short.

The mean is the default summary and it is frequently the wrong one. Knowing when — and what to reach for instead — is most of applied statistics in this job.

The hard part of writing this was cutting it. The first version covered every case; the useful version covers the case you hit on a Tuesday and names the rest in a sentence. Completeness is a property of reference material, not of teaching material, and confusing the two produces something nobody finishes.

It sits in the Statistics you will actually use module, and the exercise runs against ride-hail-trips. That pairing is deliberate: the dataset was built with the trap the lesson describes already in it, so the exercise fails in the instructive way rather than the confusing one.

Ordering matters here more than in most courses. It follows The first hour with an unfamiliar dataset and leads into Uncertainty, sampling, and how much to trust a number, and reading it out of sequence mostly works but costs you the setup.

Free means free, and it also means we can rewrite it whenever it is wrong. No edition, no errata PDF, nothing to repurchase. The whole course is 36 lessons and the fixes land the day we find them.