Field note · August 16, 2024

data-job-postings.company_size: 4 values, 81.1% in the top three

4 values with 29.6% concentrated in one of them, and what that does to every chart downstream.

1 min read ·Visualization ·visualization practice

Distribution of company_size on data-job-postings, because every chart built on it inherits this shape.

  • enterprise — 1,035 rows, 29.6%
  • scaleup — 979 rows, 28.0%
  • midmarket — 825 rows, 23.6%
  • startup — 661 rows, 18.9%

The top three take 81.1% between them. With only 4 values there is no tail to worry about, which makes this a genuinely easy column to chart.

sql
select company_size,
       count(*)                                      as rows,
       round(100.0 * count(*) / sum(count(*)) over (), 1) as pct,
       round(avg(salary_min_usd)::numeric, 2)             as avg_salary_min_usd
from data_job_postings
group by 1
order by rows desc;

The second column is the one that matters. Share of rows tells you what is common; avg_salary_min_usd tells you whether the common thing is the important thing. They disagree more often than not, and a chart that shows only the first is answering the easier question.

Decide what happens to the tail before you plot it. "Other" as an explicit bucket is honest; twelve slivers is not, and neither is silently taking the top eight.