A clean A/B test compares two versions of a page that differ in exactly one element, runs both at the same time on real visitors, and gets evaluated only once enough data and enough time have passed. Skip any of those three conditions and a test produces results that look meaningful but are really just noise.
Why most A/B tests fail
The principle sounds simple: two versions, more clicks wins. In practice, many tests still fail, either stopped too early, changing several elements at once, or running with too few visitors to ever produce a reliable result. The outcome: decisions based on chance rather than real insight.
The four basic rules for a clean test
- A clear hypothesis: not "we're testing the button" but "a green button instead of blue increases the click rate because it stands out more against the background". A hypothesis can be confirmed or disproved afterwards, a vague goal cannot.
- Only one variable: change the button colour, headline or image on its own. Change several elements at once and you will not know afterwards which one made the difference.
- Enough visitors and enough time: a test needs a minimum number of conversions per variant to be statistically meaningful, and at least one full weekly cycle so day-of-week swings even out.
- Stopping rules set in advance: decide before you start how long the test runs and how many conversions are needed. Stopping midway because one variant is briefly ahead almost always overstates the result.
Common mistakes in practice
Beyond the four basic rules, tests often fail because of smaller but consequential mistakes: the test runs over a public holiday or a sale and skews the picture, the two variants are not really shown at the same time, or a "winner" gets declared even though the difference falls within normal variation. Another well-known effect is the novelty effect: a new version often performs better in the first few days simply because returning visitors notice it as unfamiliar. That effect usually fades after a week or two, one more reason not to end tests too early. Understanding what a good conversion rate actually looks like helps put results in perspective instead of celebrating every small swing as a win.
When your traffic is too low for a classic test
Most small business websites simply do not get enough visitors for classic A/B testing with statistical significance. That does not mean optimisation is impossible, only that other tools make more sense:
- Sequential testing: run variant A for four weeks, then variant B for four weeks, and compare the results. Less clean than a true A/B test, but better than guessing.
- Qualitative methods: user interviews, session recordings and heatmaps often show where visitors get stuck after only a handful of sessions, without needing hundreds of conversions.
- Bigger, bolder changes: with little traffic, a fundamental redesign is often more useful than a small button-colour test, since its effect is visible even without strict statistics.
What comes after the test
A test does not end with the result. Document the hypothesis, the setup and the outcome, even when a variant makes no difference: a well-documented null result stops the same idea from getting tested again a year later. And a winning test is not a resting point, it is the starting point for the next hypothesis.
For ways to improve a landing page overall, even without a large test volume, see the article on landing page optimisation. For a structured approach beyond a single page, see the CRO consulting page.
Want to put this to work for your business? We review your website for free in classic and AI search and show the biggest levers.