What Is an A/B Test and How to Run One
What is an AB test and how is it run correctly? What to test, how long to wait, the typical statistical mistakes and realistic alternatives for small traffic.

The classic argument of marketing meetings: "the red button is better" — "no, the blue." Two hours, zero result. The A/B test is that argument's end: both variants get shown to real users, and the number makes the decision. When opinions are equal, the data referees.
The short answer to "what is an AB test": splitting the audience into two random groups, showing each a different variant and comparing the result. It sounds simple; in practice most tests are built wrongly and read wrongly. This article is the rules of building and reading them correctly.
What to test: the impact hierarchy
A test's value hangs on the variable's weight. The high-impact variables: the offer itself (the price model, the guarantee, the bonus), the headline/promise, the landing structure, the CTA's content, the ad creative. The low-impact ones: the button colour, small text details; those are for "having tested." The rule: on small traffic test only the big variables; your statistics will not stretch to seeing the micro-variables' difference.
The five conditions of a correct build
- One test, one variable: if the headline + image + price changed together, why did the winner win; unknown. (Multivariate tests exist, but they are a big-traffic game.)
- The same time window: not A this week and B next week; a parallel split. A week's difference is a season/campaign difference.
- Random assignment: the tools (the ad platforms' built-in A/B, the landing tools) do this themselves; hand assignment is a bias source.
- A pre-defined metric: by what will the winner be chosen; clicks, enquiries, sales? Choosing the metric after the test ends is result "hunting."
- Sufficient volume: the section below; the most-violated condition.
How long to wait: statistical patience
Two rules practically suffice. First: the minimum volume; do not believe differences before seeing at least ~100 conversion events per group (a rough orientation; small differences need more). Second: the minimum duration; at least 1–2 full weeks (the weekday/weekend behaviour difference). The most common error is the "B won!" verdict in the first 2 days: at small numbers the differences are chance's dance. The second common error is checking the test daily and stopping the moment a difference shows (peeking); that is the systematic way of declaring chance the winner. Look at the tools' "statistical confidence" percentage, but if the volume+duration conditions are unmet, that percentage deceives too.
The realistic road for small traffic
An honest admission: on a site with 500 visitors a month, the classic A/B test mathematically does not work. The alternatives: sequential comparison (2 weeks A, 2 weeks B; weak, but it catches big differences), testing at the ad level (creative/headline tests gather many events on little budget; the fastest arena for copy testing), qualitative methods (observing 5 users often says more than a 5-week test) and big changes (if you cannot see a 10% difference, test a change that will create a 50% one).
The testing culture: losing is a result too
The healthy expectation: even in good teams, most tests end with "no difference" or "it lost." That is not a loss: every result sharpens your audience model. Keep a journal: hypothesis → variants → result → learned. A six-month journal is worth more than any "best practices" list, because it is about your audience. And the winning variant is not eternal: the audience, the season, the competition change; an annual re-check of the main pages is a normal rhythm.
Frequently asked questions about A/B tests
Which free tools can one start with?
The ad platforms' built-in tests (Meta, Google Ads), the email tools' subject tests and the landing builders' split features; most are included in the plan. The need for a separate testing platform comes when traffic grows.
Can one A/B test SEO pages?
Content variants can be split across users, but showing different content to the search robot is risky territory. On the SEO side the practical road is sequential change + GSC measurement (a title change → the CTR comparison).
The test ended, the difference is 5%; do I apply it?
If the volume suffices and the metric is a business metric, yes; a 5% conversion lift is a big number over a year. If the volume is small, the 5% may be noise; extend the test or seek a variant that will enlarge the difference.
How many tests can run at once?
Parallel tests on different pages/channels are normal; overlapping tests on the same flow pollute the results. The workable rhythm for a small team: 1–2 tests at a time, but without stopping.
Professional support
Want to tie the decisions to a data system?
For diagnostics, priorities and implementation architecture, see the Revenue & Conversion Systems service.
Sources and further reading
Where to verify the source
The official documents for the platforms' built-in test tools:
Continuing the topic
The places where the test results go to work:
- The landing page's construction
- Ad copy that sells
- GA4: the measurement foundation
- Marketing KPIs
- Other articles on this topic
The best place for your first test is the ad headline: a fast event flow, a clear metric, a cheap lesson. Write one hypothesis and keep your hands off for two weeks; the numbers will do everything else.
I'm Anar Rustamli - a strategist, entrepreneur, and AI adoption leader working at the edge of growth, technology, and human thinking. Since 2016, my work has focused on helping businesses evolve in a rapidly changing digital landscape. I design growth systems, AI-powered workflows, and strategic frameworks that align performance with purpose. I believe real growth happens when strategy, data, and human insight work together - and my mission is to help businesses adopt AI in a way that strengthens both their results and their identity.

