A/B tests: what the numbers mean
Two subjects, two designs, or a second version of an automation step — and an honest answer about whether one is really winning.

EmailAnywhere you can write an email you can write two. On a newsletter, tick “test a second subject” and set the share; inside an automation, tick it on an email step and every contact entering from then on gets one version or the other. Which version a contact gets is decided from their own record, so it never changes under them — a retry, a resend, a second pass all keep them on the same side.
Email → A/B tests lists every test with what each version got and one line saying what it means. That line is a two-proportion z-test at 95% confidence — the standard A/B test. It is deliberately hard to please: each version needs at least thirty sends before it says anything at all, and a small gap on a small list is reported as “no clear winner yet”, with roughly how many more sends it would take to call. That number is the useful one. A 40-person list cannot settle a 3% difference, and a tool that pretends otherwise costs you the decision.
What to test, in order of how much it moves: the subject line, the first sentence, the offer, then the design. One change at a time — two changes and a win tells you nothing about which change won.
Sign-up forms test too. On a form (Email → Sign-up forms) the A/B test tab gives you a second set of words — the line above, the headline, the line under it and the button — while the design, the targeting and the fields stay shared, for the same reason: a form that differs in ten ways teaches nothing about any of them. Which version a visitor sees is decided from their own session, so a reload never changes the form under them, and the side they saw is recorded on every view and every sign-up. The verdict there compares sign-ups against views with the same test, so “37% better” means 37% more of the people who saw it signed up.
Reading the numbers honestly: opens undercount, because some mail apps block the tracking image, but they undercount both versions equally, so the comparison holds even when the absolute number does not. Clicks are exact. And do not stop a test the moment it turns green on a thin sample — that is how a coin flip becomes a strategy.