Field note · August 22, 2025
The category that eats the chart
6 values with 21.9% concentrated in one of them, and what that does to every chart downstream.
carrier on flight-delays has 6 values, and one of them is 21.9% of the data.
AA— 1,751 rows, 21.9%DL— 1,684 rows, 21.1%WN— 1,652 rows, 20.7%UA— 1,422 rows, 17.8%B6— 766 rows, 9.6%
The top three take 63.6% between them. With only 6 values there is no tail to worry about, which makes this a genuinely easy column to chart.
select carrier,
count(*) as rows,
round(100.0 * count(*) / sum(count(*)) over (), 1) as pct,
round(avg(dep_delay_min)::numeric, 2) as avg_dep_delay_min
from flight_delays
group by 1
order by rows desc;The second column is the one that matters. Share of rows tells you what is common; avg_dep_delay_min tells you whether the common thing is the important thing. They disagree more often than not, and a chart that shows only the first is answering the easier question.
Decide what happens to the tail before you plot it. "Other" as an explicit bucket is honest; twelve slivers is not, and neither is silently taking the top eight.