Growth

A/B Testing

By Jake Luo · Published Sep 6, 2026

A/B testing is running two versions of the same page, email or flow at the same time, splitting visitors between them at random, and keeping whichever version measurably wins. The randomisation and the simultaneity are the method: they are what stop a difference in the audience, the day or the season being mistaken for a difference in the design. Its unstated requirement is volume — a test only answers anything once enough people have converted inside each arm, which is why most early-stage sites cannot honestly run one.

What the split is actually controlling for

The instinct behind an A/B test is to find out which version is better, but that is not quite what the split does. Shipping a new page on Tuesday and comparing the fortnight after with the fortnight before is also a comparison — it just has a launch, a newsletter, a public holiday, a competitor's announcement and a ranking change mixed into it. The randomised split exists to strip those out. Both versions run in the same window, in front of the same kind of traffic, and the only thing that differs between the two groups is the one thing you changed.

That is also why naming a single primary metric before you start matters more than the tooling does. If you decide afterwards which number to look at, you will find a winner in almost any dataset, because with enough metrics on the table one of them is always up. Choosing in advance is what turns the exercise from a search for good news into a test.

The four things a test needs before it can tell you anything

A test missing any one of these does not give you a weaker answer. It gives you a confident one that happens to be wrong, which is worse, because you will act on it.

  • One primary metric, chosen in advance — signups, say, and not whichever of signups, clicks and time-on-page turns out to have moved.
  • Enough conversions in each arm — hundreds, not dozens. Visitors are not the constraint, conversions are, so a page with plenty of traffic and a thin conversion rate is still short of data.
  • A stopping rule written down first — a date or a conversion count. Checking daily and stopping the moment it looks good is how noise gets promoted to a finding.
  • A change big enough to move the number — the smaller the effect you are hunting, the larger the sample you need, and a button-colour change on a small site is asking for a resolution the traffic cannot supply.

The arithmetic behind the second point is what ends most plans. Detecting an improvement from roughly three percent to roughly four — real and well worth having — takes something on the order of several thousand visitors per arm before the result stops being a coin flip. Halve the size of the effect you are looking for and the sample you need roughly quadruples.

What our own funnel says about this

We will use our own numbers, because they are the ordinary case rather than the exception. In one recent week, 251 people reached the logged-in dashboard, 39 of them opened the plans page, and a single-digit number went on to the checkout. Those are the volumes a real early-stage product is working with, and they are not unusually bad.

Split that plans page down the middle and each version gets four or five people a week. Any difference you saw would be a difference of one or two individual decisions, which is indistinguishable from who happened to visit that week. Waiting for it to resolve means waiting months, and over those months the product, the pricing and the traffic mix all change underneath the test — so the thing you finally measure is not the thing you set out to measure.

So we make those calls by judgement, and we say that is what we are doing. The "Recommended for you" badge on our signup chooser was a deliberate decision to show one option to everybody rather than a split test, precisely because at our volume a split would have produced a number we could not have trusted. Saying so is the honest version. Describing an untested change as data-driven is the dishonest one, and it is common enough that you should assume it of most claims you read.

What to use instead until the traffic arrives

The alternatives are less satisfying and a great deal more useful at this size. Change one thing at a time and keep a dated log of what changed, so that a before-and-after at least has an edge to it. Watch what people actually do — session recordings and product analytics will show you a form field being abandoned or a button never being reached, and neither of those needs statistical power to be believed. Then talk to the people who dropped out: five conversations will out-earn an underpowered test on the same page, and they tell you why rather than only which.

Spend the effort on defects rather than preferences, too. A checkout nobody can find, a page that shows nothing for twenty seconds, a form that silently discards what was typed into it — these are not questions about which version people prefer, and testing them is a category error. Fix them, and come back to testing once the volume can carry it. The wider loop this sits inside is conversion rate optimization; when you do have the traffic, the open-source tooling looks like GrowthBook.

FAQ

What is A/B testing?
A/B testing is showing two versions of the same page, email or flow at the same time, assigning each visitor to one of them at random, and keeping whichever version performs measurably better on a metric you chose before you started. Running both simultaneously and assigning at random is the point: it removes the differences in timing and audience that make a simple before-and-after comparison unreliable.
How much traffic do I need to run an A/B test?
Think in conversions per arm, not visitors. As a working rule you want hundreds of conversions in each version before a result means much, which for a typical signup page implies several thousand visitors per arm to detect a lift of about one percentage point. Smaller effects need dramatically more: halving the size of the difference you are hunting roughly quadruples the sample.
How long should an A/B test run?
Set the stopping point before you start, as a date or a conversion count, and hold to it. Run for whole weeks rather than part-weeks, since weekday and weekend traffic behave differently. The failure mode to avoid is checking daily and stopping the moment one version is ahead — early leads reverse constantly, and a test stopped at its most flattering moment reliably reports an effect that is not there.
Can I A/B test a page that gets a few hundred visitors a month?
Not usefully. A few hundred visitors will produce a handful of conversions, split across two arms, and any gap between them is well inside the range of ordinary chance. You would either wait many months, over which the rest of the business changes and invalidates the comparison, or stop early and act on noise. At that size, sequential changes with a dated log, session recordings and conversations with people who dropped out are the instruments that actually work.
Is A/B testing the same as conversion rate optimization?
No. Conversion rate optimization is the whole loop — finding where people drop off, forming a hypothesis about why, changing something and checking whether it helped. A/B testing is one way of doing the last step, and it is only available to you once the volume supports it. Small sites still run the loop; they simply validate with a careful before-and-after and qualitative evidence instead of a split.
Related terms
Conversion Rate Optimization (CRO)Activation RateCohort analysisNorth Star Metric

An AI growth team that runs this for you

AgentCeres is a managed AI marketing team — you approve what ships. 14-day free trial, from $39/month.

Start free trialBrowse the glossary