Why it matters
Small samples produce dramatic-looking swings all the time. Ship decisions on those and you're steering the business on noise. Understanding significance, and its limits, is the difference between learning from data and being fooled by it slowly, expensively, and with great confidence.
How it works
Before the test, you fix the decision metric and the sample size. After the full sample, the observed difference is compared against the spread luck alone would create; clear it and the result is "significant". Peeking early, adding variants and swapping metrics all quietly break the guarantee.
What to do about it
Adopt one rule: decide before the data arrives (metric, sample, and what you'll do with either outcome). If a result would change a big decision, re-run it once before you re-organise around it.

