Field note · October 22, 2024
The category that eats the chart
4 values with 39.1% concentrated in one of them, and what that does to every chart downstream.
channel on retail-orders has 4 values, and one of them is 39.1% of the data.
web— 2,734 rows, 39.1%ios— 1,729 rows, 24.7%android— 1,386 rows, 19.8%marketplace— 1,151 rows, 16.4%
The top three take 83.6% between them. With only 4 values there is no tail to worry about, which makes this a genuinely easy column to chart.
select channel,
count(*) as rows,
round(100.0 * count(*) / sum(count(*)) over (), 1) as pct,
round(avg(unit_price_usd)::numeric, 2) as avg_unit_price_usd
from retail_orders
group by 1
order by rows desc;The second column is the one that matters. Share of rows tells you what is common; avg_unit_price_usd tells you whether the common thing is the important thing. They disagree more often than not, and a chart that shows only the first is answering the easier question.
Decide what happens to the tail before you plot it. "Other" as an explicit bucket is honest; twelve slivers is not, and neither is silently taking the top eight.