Field note · June 12, 2026
Deduplicate by rank, in short
Keep the best row per key with a window function instead of a self-join or a GROUP BY that loses columns.
Someone asked why Deduplicate by rank is written the way it is. Fair question.
Keep the best row per key with a window function instead of a self-join or a GROUP BY that loses columns.
What makes it a pattern rather than a tip is that the wrong version is the one you write naturally. It reads correctly, it runs, and it returns something. The failure is in the result, not in the execution — which means the only defence is recognising the shape before you are in it.
Two datasets on this site have the shape built in: dirty-customers and saas-subscriptions. Both are small enough to run the broken version, see the number, then run the corrected one and see it change.
It lives under shaping because the fix is in how the rows are arranged, not in the arithmetic. Almost every "the number is wrong" report that turns out to be real lands here.
The long-form treatment is in the course (window-functions, sql-interview-patterns); the pattern page is the version to read at 4pm with a query open.
Whether you use our version matters much less than having a version you did not re-derive under time pressure. That is what a pattern library is for.