Field note · February 10, 2024
Five minutes with prior_therapy in clinical-trial
900 rows, 2 distinct values, and the one fact about prior_therapy that changes how you query it.
Someone asked what is in prior_therapy on clinical-trial, and the honest answer took one query.
900 rows, no nulls, 2 distinct values. Whether the subject had prior therapy.
It is true in 38.0% of rows — 342 of 900.
That imbalance is the single most important fact about the column. It sets the baseline any model has to beat, it decides whether a per-segment breakdown will have enough rows in the minority class to say anything, and it determines how wide the confidence interval on any rate computed from it will be.
select
count(*) as rows,
count(prior_therapy) as present,
count(*) - count(prior_therapy) as nulls,
count(distinct prior_therapy) as distinct_values
from clinical_trial;A near-balanced flag, which is more pleasant to work with than most. Confirm the balance rather than assuming it — it is the exception, not the rule.
Where this bites: slicing by a dimension with 4 levels leaves roughly 86 true rows per slice on average. That is thin enough that the noisiest segment will look like the most extreme one, every time, and someone will read the ranking as a finding.
The grain is one row per enrolled subject, which is the context every one of those numbers depends on. None of them survive a change of grain, which is why "profile the column" and "profile the table" are the same job. Full schema, and the CSV, on the dataset page.