Field note · October 11, 2024
The overall average hides 3 different numbers
Real group means for one column across 3 segments, and what the pooled average conceals.
site splits sensor-telemetry into 3 groups. Here is what humidity_pct looks like inside each.
SLT— mean 40.95, median 41.00 (2,822 rows)NOR— mean 40.90, median 40.90 (2,823 rows)RIV— mean 40.89, median 40.90 (2,355 rows)
Top to bottom that is 40.95 against 40.89, a spread of 0.2%. The pooled average is 40.91.
select site,
count(*) as rows,
round(avg(humidity_pct)::numeric, 2) as mean,
percentile_cont(0.5) within group (order by humidity_pct) as median
from sensor_telemetry
group by 1
order by mean desc;The groups are close enough that the pooled average is a fair summary. That is worth confirming rather than assuming: the check costs one query and the failure mode is invisible.
Notice the mean and median columns disagree only slightly here. Always compute both in the group-by. The comparison between them per segment is free and tells you whether you are looking at a level difference or a tail difference.
This is the setup for Simpson's paradox — the case where every segment moves one way and the total moves the other.