Field note · April 9, 2024

Course note — The first hour with an unfamiliar dataset

A repeatable routine for finding out what you are actually holding, before you compute anything anyone will act on.

1 min read ·Data quality ·practice

Revised The first hour with an unfamiliar dataset this week. The change was small and the reason was not.

A repeatable routine for finding out what you are actually holding, before you compute anything anyone will act on.

The first draft explained the mechanism and stopped. What was missing was the failure: the thing that happens when you get it wrong, described concretely enough that someone recognises it later. A lesson that only teaches the correct version leaves you unable to spot the incorrect one, which is the situation you will actually be in.

It sits in the Data quality and trust module, and the exercise runs against dirty-customers. That pairing is deliberate: the dataset was built with the trap the lesson describes already in it, so the exercise fails in the instructive way rather than the confusing one.

Ordering matters here more than in most courses. It follows The anatomy of a silent failure and leads into Describing a column without lying, and reading it out of sequence mostly works but costs you the setup.

Free means free, and it also means we can rewrite it whenever it is wrong. No edition, no errata PDF, nothing to repurchase. The whole course is 36 lessons and the fixes land the day we find them.