Field note · July 6, 2026

Cutting a lesson down — The first hour with an unfamiliar dataset

A repeatable routine for finding out what you are actually holding, before you compute anything anyone will act on.

1 min read ·Data quality ·practice

A note from writing The first hour with an unfamiliar dataset, which took three passes to get to something short.

A repeatable routine for finding out what you are actually holding, before you compute anything anyone will act on.

The hard part of writing this was cutting it. The first version covered every case; the useful version covers the case you hit on a Tuesday and names the rest in a sentence. Completeness is a property of reference material, not of teaching material, and confusing the two produces something nobody finishes.

It sits in the Data quality and trust module, and the exercise runs against dirty-customers. That pairing is deliberate: the dataset was built with the trap the lesson describes already in it, so the exercise fails in the instructive way rather than the confusing one.

Ordering matters here more than in most courses. It follows The anatomy of a silent failure and leads into Describing a column without lying, and reading it out of sequence mostly works but costs you the setup.

Free means free, and it also means we can rewrite it whenever it is wrong. No edition, no errata PDF, nothing to repurchase. The whole course is 36 lessons and the fixes land the day we find them.