Field note · January 1, 2026

flight_date runs 2023-01-01 to 2023-12-31

365 populated days across a 365-day span, and what that allows you to compute.

1 min read ·Data engineering ·engineering quality

Coverage note for flight-delays, which is the thing to check first and the thing everyone checks last.

2023-01-01 to 2023-12-31. That is 365 calendar days, of which 365 have at least one row — so the series is dense. Average 21.9 rows per populated day.

sql
select date_trunc('day', flight_date) as day, count(*) as rows
from flight_delays
group by 1
order by 1;

A dense series means a plain group by day is safe. That is worth confirming rather than assuming — the moment a day drops out, every window function that counts rows instead of days starts comparing the wrong pair, and a seven-day lag silently becomes an eight-day one.

The other number worth having before you start: 21.9 rows per populated day tells you what granularity the data can actually support. Aggregating to something finer than the data is dense enough to fill produces a chart of noise, and there is no warning — the query returns, the line is drawn, and the wobble gets interpreted.

Because this is a date rather than a timestamp, daily aggregates are unambiguous — no time zone, no truncation, no boundary argument. That is a small thing that removes a whole category of bug, and it is worth preferring a date column whenever the time component is not genuinely used.

A date range is a fact about the data, not about the question. Filters outside it return empty results, filters half-inside it return partial ones, and neither raises anything. Check the range, then write the filter — and use a half-open interval when you do, because between on a timestamp includes exactly one instant of the final day.