Field note · December 19, 2023

Year one

What exists, what does not, and what we got wrong.

1 min read ·Analytics practice ·practice

A year ago this was a group chat. Where it actually stands:

Written: the foundations, pipelines, warehousing and SQL modules. Fourteen lessons. Considerably fewer than we planned and considerably better than the first drafts.

Built: the dataset generator, which turned out to be the best decision we made. Synthetic data means we can put the lesson in the data — informative missingness, censored durations, a Simpson's paradox — and ship it under CC0 with no licence trap.

Not built: most of the tools. We have a chart builder best described as "structurally sound and aesthetically unfinished".

Wrong about: how long writing takes. Every lesson has taken about three times the estimate, mostly in the rewriting. The first draft of a lesson takes two hours. Making it good takes six more.

Also wrong about: what people would want. We assumed the interest would be in the statistics. The questions we actually get are overwhelmingly about pipelines and data quality — how do I know it ran, how do I know it is right, what do I do when it is not.

That has changed the plan. Quality is now its own module rather than a section, and the deliberately broken dataset moved up the priority list.

Right about: free. We have had two people offer to pay and one ask where the "real" course is. There is no real course. This is it.

Next year: statistics, experiments, modelling, visualization, production, and the tools. Current estimate: six weeks.