Field note · December 19, 2023
Year one
What exists, what does not, and what we got wrong.
A year ago this was a group chat. Where it actually stands:
Written: the foundations, pipelines, warehousing and SQL modules. Fourteen lessons. Considerably fewer than we planned and considerably better than the first drafts.
Built: the dataset generator, which turned out to be the best decision we made. Synthetic data means we can put the lesson in the data — informative missingness, censored durations, a Simpson's paradox — and ship it under CC0 with no licence trap.
Not built: most of the tools. We have a chart builder best described as "structurally sound and aesthetically unfinished".
Wrong about: how long writing takes. Every lesson has taken about three times the estimate, mostly in the rewriting. The first draft of a lesson takes two hours. Making it good takes six more.
Also wrong about: what people would want. We assumed the interest would be in the statistics. The questions we actually get are overwhelmingly about pipelines and data quality — how do I know it ran, how do I know it is right, what do I do when it is not.
That has changed the plan. Quality is now its own module rather than a section, and the deliberately broken dataset moved up the priority list.
Right about: free. We have had two people offer to pay and one ask where the "real" course is. There is no real course. This is it.
Next year: statistics, experiments, modelling, visualization, production, and the tools. Current estimate: six weeks.