Field note · April 20, 2026

fare_usd: mean 22.16, median 14.29

A 55.1% gap between mean and median on one column, and which of the two answers the question you were asked.

1 min read ·Statistics ·statistics practice

A number worth keeping: on ride-hail-trips, fare_usd has a mean of 22.16 and a median of 14.29.

The mean is 22.16. The median is 14.29. That is a gap of 55.1%, and it is not noise — the p90 is 44.54 and the max is 270.61, so the top of the distribution is dragging the average somewhere a minority of rows actually live.

sql
select
  avg(fare_usd)                                                     as mean,
  percentile_cont(0.5) within group (order by fare_usd)             as median,
  percentile_cont(0.9) within group (order by fare_usd)             as p90,
  max(fare_usd)                                                     as max
from ride_hail_trips;

The practical consequence is that "average fare usd" answers a question nobody asked. If someone wants to know what a typical row looks like, the median answers it. If someone is forecasting a total, the mean is the right tool and the skew does not matter, because the mean times the count is the total.

So the rule is not "never use the mean". It is: name which question you are answering, then pick the statistic that answers it. A report that shows 22.16 with no median next to it has quietly decided you meant the first question.

There is a longer version of this in Describing a column without lying, and the pattern is the short one.