Field note · June 3, 2026
Zero is not a small number
51.3% of one column is exactly zero. Whether that is a measurement or an absence changes every aggregate built on it.
A recurring mistake, this time on discount_pct in retail-orders: treating zero as just another value on the low end.
Zero in a numeric column almost always means one of two things, and they need different handling. Either the thing was measured and came out zero, or the thing did not happen and zero is standing in for absence. 51.3% of this column is zero, which is far too much to be an accident of measurement.
Averaging across both meanings gives you a number that describes neither group. The mean of discount_pct including zeros is 0.10. Excluding them it is roughly 0.20 — a different number about a different population.
select
count(*) filter (where discount_pct = 0) as zero_rows,
count(*) filter (where discount_pct > 0) as positive_rows,
avg(discount_pct) as mean_all,
avg(discount_pct) filter (where discount_pct > 0) as mean_positive
from retail_orders;The reporting version is usually two numbers rather than one: the rate at which the thing happens, and the size when it does. Those move independently, and a single average hides which one changed. A drop in the combined mean could be fewer events or smaller events, and you cannot tell them apart after the fact.
Same shape as the null problem, except nothing warns you, because zero is a perfectly valid number and every aggregate happily includes it.