Field note · October 15, 2025
r = 0.10 between sched_dep_hour and arr_delay_min
A measured correlation of 0.10 between two columns, what its square says, and the diagnostic that costs one line.
Measured on flight-delays: the Pearson correlation between sched_dep_hour and arr_delay_min over 8,000 rows is 0.10.
That is essentially none. Squared, it says the linear relationship accounts for 1.1% of the variance in either column. Which is to say: nothing you would build anything on.
df[["sched_dep_hour", "arr_delay_min"]].corr(method="pearson")
# also worth running:
df[["sched_dep_hour", "arr_delay_min"]].corr(method="spearman")Run Spearman next to Pearson every time. They agree when the relationship is roughly linear and diverge when it is monotone but curved — and the gap between them is a free diagnostic that costs one line.
A correlation this small is compatible with a real relationship that is not linear, with a real relationship confined to one segment, and with nothing at all. It does not distinguish between them, and neither does a bigger sample.
Look at the scatter before quoting r. Anscombe's quartet is four datasets with identical correlation coefficients and nothing else in common, and the correlation explorer will draw this pair so you can see which case you are in.
Longer treatment in Correlation, confounding, and Simpson's paradox.