Field note · December 22, 2024
The overall average hides 2 different numbers
Real group means for one column across 2 segments, and what the pooled average conceals.
arm splits clinical-trial into 2 groups. Here is what age looks like inside each.
placebo— mean 54, median 53 (450 rows)treatment— mean 53, median 53 (450 rows)
Top to bottom that is 54 against 53, a spread of 1.2%. The pooled average is 53.
select arm,
count(*) as rows,
round(avg(age)::numeric, 2) as mean,
percentile_cont(0.5) within group (order by age) as median
from clinical_trial
group by 1
order by mean desc;The groups are close enough that the pooled average is a fair summary. That is worth confirming rather than assuming: the check costs one query and the failure mode is invisible.
Notice the mean and median columns disagree only slightly here. Always compute both in the group-by. The comparison between them per segment is free and tells you whether you are looking at a level difference or a tail difference.
This is the setup for Simpson's paradox — the case where every segment moves one way and the total moves the other.