Real frequentist math with no library: a two-proportion z-test for the p-value, Wilson score intervals for each arm (reliable at low counts where the normal approximation misleads), and the standard sample-size formula for a target relative lift at 80% power. Runs entirely in the browser.
5.00%
95% CI 4.42% – 5.65%
6.22%
95% CI 5.57% – 6.94%
Relative lift
+24.4%
Absolute change
+1.22 pts
z-score
2.602
The two confidence intervals still overlap— the data can't yet rule out that both arms convert identically.
To detect a 10% relative lift off a 5.00% base at 95% confidence and 80% power you need 31,233 visitors per arm — about 26,443 more each.
Most 'winning' tests we're asked to review were called early. This is the check we run first: a lift that looks decisive at 200 conversions often has confidence intervals wide enough to contain no difference at all. The math is standard — the discipline is refusing to ship until the intervals separate.