Field note · February 14, 2024

The grain uniqueness test, in short

One test per model, on the key that defines its grain. The cheapest quality control available and the one that catches the worst bugs.

1 min read ·Data quality ·quality

Added a note to The grain uniqueness test today, which is a good excuse to say the short version here.

One test per model, on the key that defines its grain. The cheapest quality control available and the one that catches the worst bugs.

What makes it a pattern rather than a tip is that the wrong version is the one you write naturally. It reads correctly, it runs, and it returns something. The failure is in the result, not in the execution — which means the only defence is recognising the shape before you are in it.

Two datasets on this site have the shape built in: retail-orders and support-tickets. Both are small enough to run the broken version, see the number, then run the corrected one and see it change.

It lives under moving because it is about data in transit: reruns, backfills, and the assumption that yesterday only ever arrives once.

The long-form treatment is in the course (testing-data-like-code, rows-grain-shape); the pattern page is the version to read at 4pm with a query open.

The pattern page has the version that holds and the version that looks right and is not, side by side. Reading them together is the point; either one alone is just code.