Field note · November 8, 2025
The category that eats the chart
3 values with 62.3% concentrated in one of them, and what that does to every chart downstream.
payment_type on ride-hail-trips has 3 values, and one of them is 62.3% of the data.
card— 4,983 rows, 62.3%wallet— 2,125 rows, 26.6%cash— 892 rows, 11.2%
The top three take 100.0% between them. With only 3 values there is no tail to worry about, which makes this a genuinely easy column to chart.
select payment_type,
count(*) as rows,
round(100.0 * count(*) / sum(count(*)) over (), 1) as pct,
round(avg(fare_usd)::numeric, 2) as avg_fare_usd
from ride_hail_trips
group by 1
order by rows desc;The second column is the one that matters. Share of rows tells you what is common; avg_fare_usd tells you whether the common thing is the important thing. They disagree more often than not, and a chart that shows only the first is answering the easier question.
Decide what happens to the tail before you plot it. "Other" as an explicit bucket is honest; twelve slivers is not, and neither is silently taking the top eight.