Field note · December 8, 2025
Five minutes with account_id in saas-subscriptions
2,500 rows, all distinct, and the one fact about account_id that changes how you query it.
Someone asked what is in account_id on saas-subscriptions, and the honest answer took one query.
2,500 rows, no nulls, 2,500 distinct values. Account identifier.
2,500 distinct values across 2,500 rows, so it is unique and a candidate key.
Uniqueness is a property of the data you have, not a guarantee about the data you will get. Nothing in the file enforces it; a re-export with an extra row, a merged upstream table, or a backfill that runs twice all break it, and none of them raise an error.
select
count(*) as rows,
count(account_id) as present,
count(*) - count(account_id) as nulls,
count(distinct account_id) as distinct_values
from saas_subscriptions;Unique today is not the same as declared unique. If nothing tests it, the day it stops being unique is the day a join starts fanning out quietly.
Where this bites: the failure is silent in both directions. A duplicated key produces too many rows and a total that is too high; a missing one produces too few and a total that is too low. Neither errors, and both look like a real change to whoever reads the number.
The grain is one row per account, which is the context every one of those numbers depends on. None of them survive a change of grain, which is why "profile the column" and "profile the table" are the same job. Full schema, and the CSV, on the dataset page.