Field note · February 4, 2026

The time coverage of city-air-quality

1,370 populated days across a 1,370-day span, and what that allows you to compute.

1 min read ·Data engineering ·engineering quality

Before any time series work on city-air-quality: what does reading_date actually cover?

2023-01-01 to 2026-10-01. That is 1,370 calendar days, of which 1,370 have at least one row — so the series is dense. Average 4.0 rows per populated day.

sql
select date_trunc('day', reading_date) as day, count(*) as rows
from city_air_quality
group by 1
order by 1;

A dense series means a plain group by day is safe. That is worth confirming rather than assuming — the moment a day drops out, every window function that counts rows instead of days starts comparing the wrong pair, and a seven-day lag silently becomes an eight-day one.

The other number worth having before you start: 4.0 rows per populated day tells you what granularity the data can actually support. Aggregating to something finer than the data is dense enough to fill produces a chart of noise, and there is no warning — the query returns, the line is drawn, and the wobble gets interpreted.

Because this is a date rather than a timestamp, daily aggregates are unambiguous — no time zone, no truncation, no boundary argument. That is a small thing that removes a whole category of bug, and it is worth preferring a date column whenever the time component is not genuinely used.

A date range is a fact about the data, not about the question. Filters outside it return empty results, filters half-inside it return partial ones, and neither raises anything. Check the range, then write the filter — and use a half-open interval when you do, because between on a timestamp includes exactly one instant of the final day.