Why it matters
Teams make hundreds of small design and copy decisions a year. Testing turns the important ones into knowledge you can bank, but only if the process is disciplined, because a false "winner" doesn't just waste the test, it hardens a wrong belief into the playbook.
How it works
Traffic is split at random, each group sees one variant, and the difference in a pre-chosen metric is judged against what chance alone would produce. The discipline is all in the setup: one primary metric, a sample size fixed in advance, and no calling it early because the dashboard turned green.
What to do about it
Test few things, properly, on pages with enough traffic to decide. If your traffic can't settle a test in a few weeks, prefer bigger swings and qualitative feedback over statistical theatre.

