Archive
2023
249 field notes published in 2023.
December 31
- A essentially none correlation, and what it is not Statistics
- Is line_no + category the grain of retail-orders? Analytics engineering
- Reading temp_c before trusting it in city-air-quality Data quality
- Course note — Turning a request into a question Analytics practice
- baseline_score by site: a 4% spread Analytics practice
- r = -0.08 between discount_pct and revenue_usd Statistics
- What pm10 actually contains in city-air-quality Data quality
- Course note — What a pipeline is really doing Data engineering
- Is user_id + exposed_on the grain of ab-test-checkout? Analytics engineering
- What reopened actually contains in support-tickets Data quality
- The overall average hides 5 different numbers Analytics practice
- r = 0.71 between unit_price_usd and revenue_usd Statistics
- Year one Analytics practice
- Reader question — Frame the problem, then beat a stupid baseline Machine learning
- Reading queue before trusting it in support-tickets Data quality
- Is movie_id + genre the grain of movie-ratings? Analytics engineering
- The overall average hides 6 different numbers Analytics practice
- Five minutes with distance_km in ride-hail-trips Data quality
- r = 0.45 between pm10 and no2_ppb Statistics
- Course note — Validation that matches deployment Machine learning
- Reading exposed_on before trusting it in ab-test-checkout Data quality
- Is resolved_at + queue the grain of support-tickets? Analytics engineering
- r = 0.98 between seats and mrr_usd Statistics
- Reading no2_ppb before trusting it in city-air-quality Data quality
- Counting rows is not testing grain Analytics engineering
- sessions by device: a 3% spread Analytics practice
- Reading signed_up_on before trusting it in saas-subscriptions Data quality
- Reader question — Owning something that runs without you Data platform
- The overall average hides 6 different numbers Analytics practice
- What reading_at actually contains in sensor-telemetry Data quality
- A moderate correlation, and what it is not Statistics
November 30
- Cutting a lesson down — What a pipeline is really doing Data engineering
- What site actually contains in clinical-trial Data quality
- Ten months in, what writing the course has taught us Analytics engineering
- Is user_id + variant the grain of ab-test-checkout? Analytics engineering
- The overall average hides 4 different numbers Analytics practice
- r = 0.25 between year and gdp_per_capita_usd Statistics
- What week12_score actually contains in clinical-trial Data quality
- Cutting a lesson down — Communicating a result to someone who will not read your notebook Visualization
- Counting rows is not testing grain Analytics engineering
- What converted actually contains in ab-test-checkout Data quality
- mrr_usd by industry: a 49% spread Analytics practice
- Is trip_id + pickup_at the grain of ride-hail-trips? Analytics engineering
- What priority actually contains in support-tickets Data quality
- A essentially none correlation, and what it is not Statistics
- Reader question — Uncertainty, sampling, and how much to trust a number Statistics
- Five minutes with sessions in ab-test-checkout Data quality
- The overall average hides 5 different numbers Analytics practice
- r = 0.98 between pm25 and pm10 Statistics
- Reading returned before trusting it in retail-orders Data quality
- humidity_pct by site: a 0% spread Analytics practice
- Cutting a lesson down — Frame the problem, then beat a stupid baseline Machine learning
- Five minutes with revenue_usd in ab-test-checkout Data quality
- Counting rows is not testing grain Analytics engineering
- temp_c by status: a 26% spread Analytics practice
- Five minutes with posted_on in data-job-postings Data quality
- r = 0.94 between gdp_per_capita_usd and co2_tonnes_per_capita Statistics
- Cutting a lesson down — When not to build a model Machine learning
- Five minutes with status in sensor-telemetry Data quality
- Is order_id + channel the grain of retail-orders? Analytics engineering
- A essentially none correlation, and what it is not Statistics
October 31
- Course note — Why your query costs what it costs Data platform
- Counting rows is not testing grain Analytics engineering
- Reading csat before trusting it in support-tickets Data quality
- salary_max_usd by country: a 379% spread Analytics practice
- A weak correlation, and what it is not Statistics
- Reader question — The anatomy of a silent failure Data quality
- Reading opened_at before trusting it in support-tickets Data quality
- Yesterday keeps changing and that is correct Data engineering
- Is line_no + country the grain of retail-orders? Analytics engineering
- bmi by sex: a 1% spread Analytics practice
- Reading sched_dep_hour before trusting it in flight-delays Data quality
- r = 1.00 between salary_min_usd and salary_max_usd Statistics
- Counting rows is not testing grain Analytics engineering
- Five minutes with ticket_id in support-tickets Data quality
- The overall average hides 4 different numbers Analytics practice
- r = -0.00 between bmi and week12_score Statistics
- Five minutes with payment_type in ride-hail-trips Data quality
- The overall average hides 6 different numbers Analytics practice
- r = 0.02 between unit_price_usd and discount_pct Statistics
- Five minutes with hour_at in grid-energy-load Data quality
- salary_min_usd by seniority: a 162% spread Analytics practice
- A moderate correlation, and what it is not Statistics
- Five minutes with pm10 in city-air-quality Data quality
- flight_no by carrier: a 3% spread Analytics practice
- A strong correlation, and what it is not Statistics
- Reading passengers before trusting it in ride-hail-trips Data quality
- The overall average hides 3 different numbers Analytics practice
- A weak correlation, and what it is not Statistics
- Reading seniority before trusting it in data-job-postings Data quality
- The overall average hides 2 different numbers Analytics practice
- A essentially none correlation, and what it is not Statistics
September 30
- What exposed_on actually contains in ab-test-checkout Data quality
- The overall average hides 4 different numbers Analytics practice
- r = 0.80 between duration_min and fare_usd Statistics
- Reading site before trusting it in sensor-telemetry Data quality
- Reader question — The first hour with an unfamiliar dataset Data quality
- r = -0.01 between bmi and baseline_score Statistics
- Five minutes with industry in saas-subscriptions Data quality
- The overall average hides 3 different numbers Analytics practice
- r = 0.01 between dep_delay_min and distance_mi Statistics
- What adverse_event actually contains in clinical-trial Data quality
- Cutting a lesson down — The five tests that cover most questions Experimentation
- DuckDB replaced a Spark cluster for one of our jobs Data platform
- The overall average hides 4 different numbers Analytics practice
- r = -0.41 between pm25 and temp_c Statistics
- Five minutes with week12_score in clinical-trial Data quality
- The overall average hides 30 different numbers Analytics practice
- Counting rows is not testing grain Analytics engineering
- Five minutes with seats_now in saas-subscriptions Data quality
- The overall average hides 4 different numbers Analytics practice
- r = -0.05 between population and gdp_per_capita_usd Statistics
- What rating actually contains in movie-ratings Data quality
- Reader question — One metric, one definition Analytics engineering
- A essentially none correlation, and what it is not Statistics
- Five minutes with ordered_at in retail-orders Data quality
- age by arm: a 1% spread Analytics practice
- r = -0.01 between sched_dep_hour and distance_mi Statistics
- Reading temp_c before trusting it in sensor-telemetry Data quality
- Counting rows is not testing grain Analytics engineering
- The overall average hides 5 different numbers Analytics practice
- Reading tip_usd before trusting it in ride-hail-trips Data quality
August 31
- r = -0.01 between flight_no and distance_mi Statistics
- The overall average hides 8 different numbers Analytics practice
- Reader question — Batch, streaming, and the honest difference Data engineering
- A essentially none correlation, and what it is not Statistics
- revenue_usd by device: a 28% spread Analytics practice
- Reading cancelled before trusting it in flight-delays Data quality
- r = -0.40 between pm10 and temp_c Statistics
- Cutting a lesson down — Validation that matches deployment Machine learning
- What gdp_per_capita_usd actually contains in world-indicators Data quality
- We p-hacked ourselves and caught it in review Experimentation
- r = -0.02 between sessions and revenue_usd Statistics
- Cutting a lesson down — Colour is an encoding, not decoration Visualization
- duration_min by borough: a 254% spread Analytics practice
- A essentially none correlation, and what it is not Statistics
- Reading o3_ppb before trusting it in city-air-quality Data quality
- unit_price_usd by channel: a 6% spread Analytics practice
- r = 0.01 between arr_delay_min and distance_mi Statistics
- Five minutes with site in sensor-telemetry Data quality
- distance_km by borough: a 331% spread Analytics practice
- Cutting a lesson down — Window functions, properly Analytics engineering
- A essentially none correlation, and what it is not Statistics
- temp_c by site: a 3% spread Analytics practice
- Five minutes with surge_multiplier in ride-hail-trips Data quality
- A essentially none correlation, and what it is not Statistics
- year by country: a wide spread Analytics practice
- Five minutes with load_mw in grid-energy-load Data quality
- A essentially none correlation, and what it is not Statistics
- The overall average hides 3 different numbers Analytics practice
- Five minutes with arr_delay_min in flight-delays Data quality
- A moderate correlation, and what it is not Statistics
- revenue_usd by category: a 1,000% spread Analytics practice
July 31
- What internet_pct actually contains in world-indicators Data quality
- Reader question — Validation that matches deployment Machine learning
- r = 0.08 between temp_c and humidity_pct Statistics
- Five minutes with unit_price_usd in retail-orders Data quality
- Is account_id + region the grain of saas-subscriptions? Analytics engineering
- Cutting a lesson down — Uncertainty, sampling, and how much to trust a number Statistics
- We deleted 34 dashboards and nobody complained Visualization
- r = 0.26 between year and life_expectancy Statistics
- Reading revenue_usd before trusting it in retail-orders Data quality
- vibration_mm_s by site: a 29% spread Analytics practice
- A essentially none correlation, and what it is not Statistics
- Reading trip_id before trusting it in ride-hail-trips Data quality
- The overall average hides 6 different numbers Analytics practice
- A strong correlation, and what it is not Statistics
- Reading status before trusting it in sensor-telemetry Data quality
- The overall average hides 30 different numbers Analytics practice
- A strong correlation, and what it is not Statistics
- Reading revenue_usd before trusting it in ab-test-checkout Data quality
- The overall average hides 10 different numbers Analytics practice
- A essentially none correlation, and what it is not Statistics
- What year actually contains in world-indicators Data quality
- Counting rows is not testing grain Analytics engineering
- Course note — Batch, streaming, and the honest difference Data engineering
- The overall average hides 3 different numbers Analytics practice
- r = 0.22 between year and co2_tonnes_per_capita Statistics
- Reading city before trusting it in data-job-postings Data quality
- duration_min by payment_type: a 3% spread Analytics practice
- Cutting a lesson down — The patterns that keep coming up Analytics engineering
- Five minutes with co2_tonnes_per_capita in world-indicators Data quality
- The overall average hides 14 different numbers Analytics practice
- Counting rows is not testing grain Analytics engineering
June 30
- Course note — Owning something that runs without you Data platform
- population by region: a 63% spread Analytics practice
- r = 0.01 between no2_ppb and temp_c Statistics
- Schema drift, on a Friday, obviously Data engineering
- Five minutes with salary_min_usd in data-job-postings Data quality
- The overall average hides 4 different numbers Analytics practice
- A weak correlation, and what it is not Statistics
- Five minutes with gdp_per_capita_usd in world-indicators Data quality
- Reader question — What a pipeline is really doing Data engineering
- The overall average hides 14 different numbers Analytics practice
- r = 0.08 between solar_mw and price_eur_mwh Statistics
- Reading co2_tonnes_per_capita before trusting it in world-indicators Data quality
- pm25 by city: a 173% spread Analytics practice
- r = 0.00 between flight_no and dep_delay_min Statistics
- What account_id actually contains in saas-subscriptions Data quality
- Reader question — Testing data like you test code Data quality
- r = -0.01 between age and baseline_score Statistics
- Five minutes with pm25 in city-air-quality Data quality
- Course note — Orchestration, and what "it runs every night" costs Data platform
- The overall average hides 8 different numbers Analytics practice
- Reading origin before trusting it in flight-delays Data quality
- A moderate correlation, and what it is not Statistics
- Reader question — Why your query costs what it costs Data platform
- The overall average hides 4 different numbers Analytics practice
- A essentially none correlation, and what it is not Statistics
- Reading priority before trusting it in support-tickets Data quality
- Course note — Correlation, confounding, and Simpson's paradox Statistics
- revenue_usd by country: a 12% spread Analytics practice
- r = -0.36 between pm25 and o3_ppb Statistics
- What csat actually contains in support-tickets Data quality
May 31
- The overall average hides 3 different numbers Analytics practice
- The meeting where the median won Statistics
- Course note — Slowly changing dimensions, and the "as of when" problem Analytics engineering
- Five minutes with reading_date in city-air-quality Data quality
- The overall average hides 4 different numbers Analytics practice
- A strong correlation, and what it is not Statistics
- Five minutes with resolved_at in support-tickets Data quality
- Reader question — Monitoring data, not just jobs Data platform
- A strong correlation, and what it is not Statistics
- Reading channel before trusting it in retail-orders Data quality
- year by region: a wide spread Analytics practice
- r = 0.11 between release_year and rating Statistics
- Five minutes with city in data-job-postings Data quality
- baseline_score by arm: a 0% spread Analytics practice
- r = 0.96 between dep_delay_min and arr_delay_min Statistics
- What solar_mw actually contains in grid-energy-load Data quality
- The overall average hides 8 different numbers Analytics practice
- r = 1.00 between seats_now and mrr_usd Statistics
- Five minutes with dep_delay_min in flight-delays Data quality
- The overall average hides 8 different numbers Analytics practice
- r = 0.46 between pm25 and no2_ppb Statistics
- What salary_min_usd actually contains in data-job-postings Data quality
- The overall average hides 6 different numbers Analytics practice
- A essentially none correlation, and what it is not Statistics
- Reading country before trusting it in world-indicators Data quality
- Reader question — Describing a column without lying Statistics
- A essentially none correlation, and what it is not Statistics
- Reading converted before trusting it in ab-test-checkout Data quality
- unit_price_usd by category: a 1,012% spread Analytics practice
- A moderate correlation, and what it is not Statistics
- What first_response_min actually contains in support-tickets Data quality
April 1
- The notebook that only ran once Machine learning
March 1
- We saved $2,100 a month by typing more Data platform
February 1
- The join that doubled revenue for six weeks Data quality
January 1
- Why we started writing this down Analytics practice