Archive

2023

249 field notes published in 2023.

December 31

  1. A essentially none correlation, and what it is not Statistics
  2. Is line_no + category the grain of retail-orders? Analytics engineering
  3. Reading temp_c before trusting it in city-air-quality Data quality
  4. Course note — Turning a request into a question Analytics practice
  5. baseline_score by site: a 4% spread Analytics practice
  6. r = -0.08 between discount_pct and revenue_usd Statistics
  7. What pm10 actually contains in city-air-quality Data quality
  8. Course note — What a pipeline is really doing Data engineering
  9. Is user_id + exposed_on the grain of ab-test-checkout? Analytics engineering
  10. What reopened actually contains in support-tickets Data quality
  11. The overall average hides 5 different numbers Analytics practice
  12. r = 0.71 between unit_price_usd and revenue_usd Statistics
  13. Year one Analytics practice
  14. Reader question — Frame the problem, then beat a stupid baseline Machine learning
  15. Reading queue before trusting it in support-tickets Data quality
  16. Is movie_id + genre the grain of movie-ratings? Analytics engineering
  17. The overall average hides 6 different numbers Analytics practice
  18. Five minutes with distance_km in ride-hail-trips Data quality
  19. r = 0.45 between pm10 and no2_ppb Statistics
  20. Course note — Validation that matches deployment Machine learning
  21. Reading exposed_on before trusting it in ab-test-checkout Data quality
  22. Is resolved_at + queue the grain of support-tickets? Analytics engineering
  23. r = 0.98 between seats and mrr_usd Statistics
  24. Reading no2_ppb before trusting it in city-air-quality Data quality
  25. Counting rows is not testing grain Analytics engineering
  26. sessions by device: a 3% spread Analytics practice
  27. Reading signed_up_on before trusting it in saas-subscriptions Data quality
  28. Reader question — Owning something that runs without you Data platform
  29. The overall average hides 6 different numbers Analytics practice
  30. What reading_at actually contains in sensor-telemetry Data quality
  31. A moderate correlation, and what it is not Statistics

November 30

  1. Cutting a lesson down — What a pipeline is really doing Data engineering
  2. What site actually contains in clinical-trial Data quality
  3. Ten months in, what writing the course has taught us Analytics engineering
  4. Is user_id + variant the grain of ab-test-checkout? Analytics engineering
  5. The overall average hides 4 different numbers Analytics practice
  6. r = 0.25 between year and gdp_per_capita_usd Statistics
  7. What week12_score actually contains in clinical-trial Data quality
  8. Cutting a lesson down — Communicating a result to someone who will not read your notebook Visualization
  9. Counting rows is not testing grain Analytics engineering
  10. What converted actually contains in ab-test-checkout Data quality
  11. mrr_usd by industry: a 49% spread Analytics practice
  12. Is trip_id + pickup_at the grain of ride-hail-trips? Analytics engineering
  13. What priority actually contains in support-tickets Data quality
  14. A essentially none correlation, and what it is not Statistics
  15. Reader question — Uncertainty, sampling, and how much to trust a number Statistics
  16. Five minutes with sessions in ab-test-checkout Data quality
  17. The overall average hides 5 different numbers Analytics practice
  18. r = 0.98 between pm25 and pm10 Statistics
  19. Reading returned before trusting it in retail-orders Data quality
  20. humidity_pct by site: a 0% spread Analytics practice
  21. Cutting a lesson down — Frame the problem, then beat a stupid baseline Machine learning
  22. Five minutes with revenue_usd in ab-test-checkout Data quality
  23. Counting rows is not testing grain Analytics engineering
  24. temp_c by status: a 26% spread Analytics practice
  25. Five minutes with posted_on in data-job-postings Data quality
  26. r = 0.94 between gdp_per_capita_usd and co2_tonnes_per_capita Statistics
  27. Cutting a lesson down — When not to build a model Machine learning
  28. Five minutes with status in sensor-telemetry Data quality
  29. Is order_id + channel the grain of retail-orders? Analytics engineering
  30. A essentially none correlation, and what it is not Statistics

October 31

  1. Course note — Why your query costs what it costs Data platform
  2. Counting rows is not testing grain Analytics engineering
  3. Reading csat before trusting it in support-tickets Data quality
  4. salary_max_usd by country: a 379% spread Analytics practice
  5. A weak correlation, and what it is not Statistics
  6. Reader question — The anatomy of a silent failure Data quality
  7. Reading opened_at before trusting it in support-tickets Data quality
  8. Yesterday keeps changing and that is correct Data engineering
  9. Is line_no + country the grain of retail-orders? Analytics engineering
  10. bmi by sex: a 1% spread Analytics practice
  11. Reading sched_dep_hour before trusting it in flight-delays Data quality
  12. r = 1.00 between salary_min_usd and salary_max_usd Statistics
  13. Counting rows is not testing grain Analytics engineering
  14. Five minutes with ticket_id in support-tickets Data quality
  15. The overall average hides 4 different numbers Analytics practice
  16. r = -0.00 between bmi and week12_score Statistics
  17. Five minutes with payment_type in ride-hail-trips Data quality
  18. The overall average hides 6 different numbers Analytics practice
  19. r = 0.02 between unit_price_usd and discount_pct Statistics
  20. Five minutes with hour_at in grid-energy-load Data quality
  21. salary_min_usd by seniority: a 162% spread Analytics practice
  22. A moderate correlation, and what it is not Statistics
  23. Five minutes with pm10 in city-air-quality Data quality
  24. flight_no by carrier: a 3% spread Analytics practice
  25. A strong correlation, and what it is not Statistics
  26. Reading passengers before trusting it in ride-hail-trips Data quality
  27. The overall average hides 3 different numbers Analytics practice
  28. A weak correlation, and what it is not Statistics
  29. Reading seniority before trusting it in data-job-postings Data quality
  30. The overall average hides 2 different numbers Analytics practice
  31. A essentially none correlation, and what it is not Statistics

September 30

  1. What exposed_on actually contains in ab-test-checkout Data quality
  2. The overall average hides 4 different numbers Analytics practice
  3. r = 0.80 between duration_min and fare_usd Statistics
  4. Reading site before trusting it in sensor-telemetry Data quality
  5. Reader question — The first hour with an unfamiliar dataset Data quality
  6. r = -0.01 between bmi and baseline_score Statistics
  7. Five minutes with industry in saas-subscriptions Data quality
  8. The overall average hides 3 different numbers Analytics practice
  9. r = 0.01 between dep_delay_min and distance_mi Statistics
  10. What adverse_event actually contains in clinical-trial Data quality
  11. Cutting a lesson down — The five tests that cover most questions Experimentation
  12. DuckDB replaced a Spark cluster for one of our jobs Data platform
  13. The overall average hides 4 different numbers Analytics practice
  14. r = -0.41 between pm25 and temp_c Statistics
  15. Five minutes with week12_score in clinical-trial Data quality
  16. The overall average hides 30 different numbers Analytics practice
  17. Counting rows is not testing grain Analytics engineering
  18. Five minutes with seats_now in saas-subscriptions Data quality
  19. The overall average hides 4 different numbers Analytics practice
  20. r = -0.05 between population and gdp_per_capita_usd Statistics
  21. What rating actually contains in movie-ratings Data quality
  22. Reader question — One metric, one definition Analytics engineering
  23. A essentially none correlation, and what it is not Statistics
  24. Five minutes with ordered_at in retail-orders Data quality
  25. age by arm: a 1% spread Analytics practice
  26. r = -0.01 between sched_dep_hour and distance_mi Statistics
  27. Reading temp_c before trusting it in sensor-telemetry Data quality
  28. Counting rows is not testing grain Analytics engineering
  29. The overall average hides 5 different numbers Analytics practice
  30. Reading tip_usd before trusting it in ride-hail-trips Data quality

August 31

  1. r = -0.01 between flight_no and distance_mi Statistics
  2. The overall average hides 8 different numbers Analytics practice
  3. Reader question — Batch, streaming, and the honest difference Data engineering
  4. A essentially none correlation, and what it is not Statistics
  5. revenue_usd by device: a 28% spread Analytics practice
  6. Reading cancelled before trusting it in flight-delays Data quality
  7. r = -0.40 between pm10 and temp_c Statistics
  8. Cutting a lesson down — Validation that matches deployment Machine learning
  9. What gdp_per_capita_usd actually contains in world-indicators Data quality
  10. We p-hacked ourselves and caught it in review Experimentation
  11. r = -0.02 between sessions and revenue_usd Statistics
  12. Cutting a lesson down — Colour is an encoding, not decoration Visualization
  13. duration_min by borough: a 254% spread Analytics practice
  14. A essentially none correlation, and what it is not Statistics
  15. Reading o3_ppb before trusting it in city-air-quality Data quality
  16. unit_price_usd by channel: a 6% spread Analytics practice
  17. r = 0.01 between arr_delay_min and distance_mi Statistics
  18. Five minutes with site in sensor-telemetry Data quality
  19. distance_km by borough: a 331% spread Analytics practice
  20. Cutting a lesson down — Window functions, properly Analytics engineering
  21. A essentially none correlation, and what it is not Statistics
  22. temp_c by site: a 3% spread Analytics practice
  23. Five minutes with surge_multiplier in ride-hail-trips Data quality
  24. A essentially none correlation, and what it is not Statistics
  25. year by country: a wide spread Analytics practice
  26. Five minutes with load_mw in grid-energy-load Data quality
  27. A essentially none correlation, and what it is not Statistics
  28. The overall average hides 3 different numbers Analytics practice
  29. Five minutes with arr_delay_min in flight-delays Data quality
  30. A moderate correlation, and what it is not Statistics
  31. revenue_usd by category: a 1,000% spread Analytics practice

July 31

  1. What internet_pct actually contains in world-indicators Data quality
  2. Reader question — Validation that matches deployment Machine learning
  3. r = 0.08 between temp_c and humidity_pct Statistics
  4. Five minutes with unit_price_usd in retail-orders Data quality
  5. Is account_id + region the grain of saas-subscriptions? Analytics engineering
  6. Cutting a lesson down — Uncertainty, sampling, and how much to trust a number Statistics
  7. We deleted 34 dashboards and nobody complained Visualization
  8. r = 0.26 between year and life_expectancy Statistics
  9. Reading revenue_usd before trusting it in retail-orders Data quality
  10. vibration_mm_s by site: a 29% spread Analytics practice
  11. A essentially none correlation, and what it is not Statistics
  12. Reading trip_id before trusting it in ride-hail-trips Data quality
  13. The overall average hides 6 different numbers Analytics practice
  14. A strong correlation, and what it is not Statistics
  15. Reading status before trusting it in sensor-telemetry Data quality
  16. The overall average hides 30 different numbers Analytics practice
  17. A strong correlation, and what it is not Statistics
  18. Reading revenue_usd before trusting it in ab-test-checkout Data quality
  19. The overall average hides 10 different numbers Analytics practice
  20. A essentially none correlation, and what it is not Statistics
  21. What year actually contains in world-indicators Data quality
  22. Counting rows is not testing grain Analytics engineering
  23. Course note — Batch, streaming, and the honest difference Data engineering
  24. The overall average hides 3 different numbers Analytics practice
  25. r = 0.22 between year and co2_tonnes_per_capita Statistics
  26. Reading city before trusting it in data-job-postings Data quality
  27. duration_min by payment_type: a 3% spread Analytics practice
  28. Cutting a lesson down — The patterns that keep coming up Analytics engineering
  29. Five minutes with co2_tonnes_per_capita in world-indicators Data quality
  30. The overall average hides 14 different numbers Analytics practice
  31. Counting rows is not testing grain Analytics engineering

June 30

  1. Course note — Owning something that runs without you Data platform
  2. population by region: a 63% spread Analytics practice
  3. r = 0.01 between no2_ppb and temp_c Statistics
  4. Schema drift, on a Friday, obviously Data engineering
  5. Five minutes with salary_min_usd in data-job-postings Data quality
  6. The overall average hides 4 different numbers Analytics practice
  7. A weak correlation, and what it is not Statistics
  8. Five minutes with gdp_per_capita_usd in world-indicators Data quality
  9. Reader question — What a pipeline is really doing Data engineering
  10. The overall average hides 14 different numbers Analytics practice
  11. r = 0.08 between solar_mw and price_eur_mwh Statistics
  12. Reading co2_tonnes_per_capita before trusting it in world-indicators Data quality
  13. pm25 by city: a 173% spread Analytics practice
  14. r = 0.00 between flight_no and dep_delay_min Statistics
  15. What account_id actually contains in saas-subscriptions Data quality
  16. Reader question — Testing data like you test code Data quality
  17. r = -0.01 between age and baseline_score Statistics
  18. Five minutes with pm25 in city-air-quality Data quality
  19. Course note — Orchestration, and what "it runs every night" costs Data platform
  20. The overall average hides 8 different numbers Analytics practice
  21. Reading origin before trusting it in flight-delays Data quality
  22. A moderate correlation, and what it is not Statistics
  23. Reader question — Why your query costs what it costs Data platform
  24. The overall average hides 4 different numbers Analytics practice
  25. A essentially none correlation, and what it is not Statistics
  26. Reading priority before trusting it in support-tickets Data quality
  27. Course note — Correlation, confounding, and Simpson's paradox Statistics
  28. revenue_usd by country: a 12% spread Analytics practice
  29. r = -0.36 between pm25 and o3_ppb Statistics
  30. What csat actually contains in support-tickets Data quality

May 31

  1. The overall average hides 3 different numbers Analytics practice
  2. The meeting where the median won Statistics
  3. Course note — Slowly changing dimensions, and the "as of when" problem Analytics engineering
  4. Five minutes with reading_date in city-air-quality Data quality
  5. The overall average hides 4 different numbers Analytics practice
  6. A strong correlation, and what it is not Statistics
  7. Five minutes with resolved_at in support-tickets Data quality
  8. Reader question — Monitoring data, not just jobs Data platform
  9. A strong correlation, and what it is not Statistics
  10. Reading channel before trusting it in retail-orders Data quality
  11. year by region: a wide spread Analytics practice
  12. r = 0.11 between release_year and rating Statistics
  13. Five minutes with city in data-job-postings Data quality
  14. baseline_score by arm: a 0% spread Analytics practice
  15. r = 0.96 between dep_delay_min and arr_delay_min Statistics
  16. What solar_mw actually contains in grid-energy-load Data quality
  17. The overall average hides 8 different numbers Analytics practice
  18. r = 1.00 between seats_now and mrr_usd Statistics
  19. Five minutes with dep_delay_min in flight-delays Data quality
  20. The overall average hides 8 different numbers Analytics practice
  21. r = 0.46 between pm25 and no2_ppb Statistics
  22. What salary_min_usd actually contains in data-job-postings Data quality
  23. The overall average hides 6 different numbers Analytics practice
  24. A essentially none correlation, and what it is not Statistics
  25. Reading country before trusting it in world-indicators Data quality
  26. Reader question — Describing a column without lying Statistics
  27. A essentially none correlation, and what it is not Statistics
  28. Reading converted before trusting it in ab-test-checkout Data quality
  29. unit_price_usd by category: a 1,012% spread Analytics practice
  30. A moderate correlation, and what it is not Statistics
  31. What first_response_min actually contains in support-tickets Data quality

April 1

  1. The notebook that only ran once Machine learning

March 1

  1. We saved $2,100 a month by typing more Data platform

February 1

  1. The join that doubled revenue for six weeks Data quality

January 1

  1. Why we started writing this down Analytics practice

Back to the blog