Field note · February 6, 2026

ride-hail-trips.borough: 5 values, 73.6% in the top three

5 values with 30.0% concentrated in one of them, and what that does to every chart downstream.

1 min read ·Visualization ·visualization practice

Distribution of borough on ride-hail-trips, because every chart built on it inherits this shape.

  • Downtown — 2,399 rows, 30.0%
  • Midtown — 2,034 rows, 25.4%
  • Uptown — 1,453 rows, 18.2%
  • Harbour — 1,149 rows, 14.4%
  • Airport — 965 rows, 12.1%

The top three take 73.6% between them. With only 5 values there is no tail to worry about, which makes this a genuinely easy column to chart.

sql
select borough,
       count(*)                                      as rows,
       round(100.0 * count(*) / sum(count(*)) over (), 1) as pct,
       round(avg(fare_usd)::numeric, 2)             as avg_fare_usd
from ride_hail_trips
group by 1
order by rows desc;

The second column is the one that matters. Share of rows tells you what is common; avg_fare_usd tells you whether the common thing is the important thing. They disagree more often than not, and a chart that shows only the first is answering the easier question.

Decide what happens to the tail before you plot it. "Other" as an explicit bucket is honest; twelve slivers is not, and neither is silently taking the top eight.