Field note · September 27, 2024
The category that eats the chart
30 values with 3.3% concentrated in one of them, and what that does to every chart downstream.
country on world-indicators has 30 values, and one of them is 3.3% of the data.
Anselm— 24 rows, 3.3%Avalonia— 24 rows, 3.3%Belmara— 24 rows, 3.3%Brasilia Nova— 24 rows, 3.3%Cairnvale— 24 rows, 3.3%
The top three take 10.0% between them. The remaining 27 share 90.0%, which is the part that gets rendered as an unreadable stack of slivers if you plot all of them.
select country,
count(*) as rows,
round(100.0 * count(*) / sum(count(*)) over (), 1) as pct,
round(avg(gdp_per_capita_usd)::numeric, 2) as avg_gdp_per_capita_usd
from world_indicators
group by 1
order by rows desc;The second column is the one that matters. Share of rows tells you what is common; avg_gdp_per_capita_usd tells you whether the common thing is the important thing. They disagree more often than not, and a chart that shows only the first is answering the easier question.
Decide what happens to the tail before you plot it. "Other" as an explicit bucket is honest; twelve slivers is not, and neither is silently taking the top eight.