25/07/202615 min read

What Is AB Testing and Why It Drives Revenue

Justine Bowman

By Justine Bowman

What Is AB Testing and Why It Drives Revenue

A/B testing is a controlled experiment that compares two versions of a page, ad, or flow against a chosen metric, so teams ship changes based on evidence instead of opinion. In South Africa, that matters because a retailer's review-feature test lifted conversion rate by 22%, which is the kind of gain that can change a week's revenue without changing traffic (Google's South African case study summary).

You're probably staring at a checkout page, a landing page, or an ad account right now and wondering whether the next tweak is a breakthrough or just expensive decoration. That's exactly where A/B testing earns its keep, because it turns “I think this looks better” into a decision you can defend with actual user behaviour.

Table of Contents

The Cost of Shipping Changes You Never Tested

A team redesigns a checkout page on Monday because the old one “felt cluttered”. By Friday, paid traffic is still flowing, the conversion rate looks softer, and nobody can prove whether the redesign caused the dip or whether the usual midweek traffic pattern did. That's the core cost of shipping on hunches, you pay for the change, pay for the traffic, then pay again in lost confidence when the result is impossible to interpret.

That's why A/B testing is best understood as a risk-reduction system, not a statistics hobby. You're not trying to win a debate in the boardroom, you're trying to avoid burning media spend on a page, ad, or funnel change that never deserved a full rollout in the first place. If you're prioritising what to test first, it helps to prioritize A/B tests for SaaS by looking at pages where small gains compound quickly, especially in the lower funnel.

A lot of teams learn this the hard way. They launch a new headline, a new hero image, or a simplified form, then declare victory after a few good-looking sessions. The trouble is that a short window can flatter the wrong version, especially when traffic is coming from different devices, ad sets, or regions.

Practical rule: if you can't explain what changed, who saw it, and which metric proved it helped, you don't have a test yet, you have a guess with dashboard support.

This is also why A/B testing shows up so often in mature CRO work. It gives founders and marketing managers a cleaner answer to a messy question, should we ship this now, keep iterating, or kill it? A disciplined test does not just reduce arguments, it protects margin by making sure paid traffic lands on experiences that have already earned the right to scale.

What A/B Testing Actually Means

A/B testing is simple in concept and very specific in execution. You take one experience, the control or A, create a changed version, the variant or B, and split similar visitors between them so you can measure which one performs better on one chosen metric. Oracle describes this as a controlled experiment, also called split testing or bucket testing, and statistically it's a two-sample hypothesis test (Oracle on A/B testing).

Think of two shop windows on the same street. One window keeps the current display, the other changes only one thing, maybe the sign, maybe the product arrangement. Passers-by are then split between those windows, and you watch which setup draws more people inside. The point isn't to admire the display, it's to learn which version changes behaviour.

A flowchart diagram illustrating the A/B testing process from experiment start to choosing the winning variation.

The terms that matter in practice

The words sound technical, but each one maps to a decision. A hypothesis is the reason you believe a change might work, for example, shortening a form should reduce friction. A primary metric is the one number that decides success, such as lead submissions, add-to-cart rate, or completed purchases.

Random assignment is what keeps the result honest, because it spreads visitors across A and B in a way that avoids obvious bias. Confidence level is the trust threshold you set before calling a result real, not a vibe you add later when the chart looks good.

A/B testing gets useful when the team agrees on the question before the launch, not after the results start leaning one way.

If you want the cleanest way to explain what is ab testing to a colleague, use this sentence. It's a controlled split between two versions, measured against one metric, so the team can ship the better performer rather than the louder opinion. That's all the structure you need before the details start to matter.

For teams that want a broader conversion lens, SelfServe conversion tips are useful because they connect testing to the bigger job of reducing friction, not just changing colours or button copy.

The Statistical Basics That Decide a Winner

A test can look like a win before it is one. A small lift on a shaky sample often disappears once more visitors arrive, because traffic patterns shift across the week and your audience does not behave the same way from one session to the next.

A flowchart diagram showing that Sample Size, Confidence Level, and MDE determine statistical significance in testing.

Sample size is the gatekeeper

Nielsen Norman Group recommends setting the baseline outcome, the minimum detectable effect, and a significance threshold usually set at 95%, then using those inputs to work out the sample size you need (NN/g on A/B testing). That matters because a test cannot separate signal from noise if too few visitors enter it.

Short runs fool teams all the time. A result that looks strong on a tiny audience can vanish when the next wave of visitors behaves differently, while the same lift on a larger sample can be worth acting on because it has room to survive variance.

Minimum detectable effect changes the conversation

A minimum detectable effect is the smallest improvement worth the effort. That is a business call, not a statistical luxury. If the change requires design time, developer time, or a real shift in workflow, the lift has to justify that work.

Decision rule: if the lift would not change what you would actually ship, the test probably is not worth running.

That is where many test plans fall apart before launch. Teams pick a page, change something, then ask later whether the gap is large enough to roll out. Better testing starts with the threshold for action, then checks whether the sample and the timeline can answer that question cleanly.

Test duration is part of validity

NN/g also recommends running the test for at least 1 to 2 weeks so you catch day-to-day behaviour changes (NN/g on A/B testing). In South African and other uneven-traffic markets, that discipline matters even more because weekday and weekend behaviour can diverge sharply. A test that ends too soon can reward coincidence, not strategy.

For a practical view on threshold setting, see what statistical significance means in practice. The useful takeaway is simple. Do not call a winner until the sample, the threshold, and the duration all point to the same answer.

For teams building a broader optimisation programme, modern conversion optimization strategy is a helpful companion because it treats test selection, confidence, and execution as one revenue system rather than isolated tasks.

Types of Experiments Worth Running

Not every question deserves the same test structure. If you only need to compare one headline, classic A/B testing is usually enough. If you want to compare several candidate headlines at once, an A/B/n test is more efficient because it compares the control against multiple variants in one run.

Pick the test type that matches the question

A multivariate test is different, because it checks how multiple elements interact with each other. That's useful when you care about combinations, like a hero image, a subheading, and a CTA that may work well individually but behave differently together. It's a stronger fit for mature teams with enough traffic to support a more fragmented sample.

A split-URL test is the right tool when the page structure itself is changing. If the team is comparing a legacy landing page against a totally different layout, or testing a new journey with its own URL, split testing keeps the experiment clean because each version lives on a separate path.

Test type Best used for What it helps answer
A/B One change on a page or ad Which version performs better on one metric
A/B/n Several candidate variants Which of several ideas deserves rollout
Multivariate Multiple interacting elements Which combination works best together
Split-URL Major layout or flow changes Which full-page experience wins

If you want a practical toolset around experiment planning, conversion rate optimisation tools are worth reviewing because the right platform makes it easier to keep the test clean, measure the right event, and avoid messy data collection.

Match the format to the business risk

A single CTA colour test is usually not worth a heavyweight multivariate setup. A pricing page overhaul, on the other hand, may deserve a split-URL test because the visitor experience changes too much for a small-page-element test to answer the question. That's the judgment call senior teams make every week.

The trap is using a more complex test because it sounds more advanced. Complex doesn't mean better. The best experiment is the smallest one that can answer the question with enough confidence to inform revenue decisions.

Real Results From eCommerce, SaaS, and Property

A useful test starts with a business question that affects revenue. For an eCommerce retailer, that could be whether social proof belongs above the fold or closer to the purchase button. Google's South African work on Mr Price is a strong example, because adding product reviews lifted the site's conversion rate by 22% and also improved product page engagement.

eCommerce is where friction shows up fast

That kind of result matters because eCommerce decisions sit close to revenue. Reviews reduce uncertainty, especially when buyers cannot touch the product or ask a sales rep. In a market with expensive paid traffic, even a modest improvement can change the economics of the whole funnel.

SaaS tests usually look different. A trial-flow change, a signup-page adjustment, or a pricing-page test often aims to cut drop-off rather than chase flashy design wins. The best SaaS experiments are usually boring to look at and valuable to the business, because they remove a reason not to convert.

Property teams need a different lens again. Lead quality matters as much as lead count, so form changes can improve enquiry quality by forcing the right amount of commitment without overwhelming the user. A shorter path is not always a better path if it floods the sales team with poor-fit requests.

An infographic showing conversion rate results from eCommerce, SaaS, and property industries using A/B testing methods.

The best tests do not just lift a number, they make the next step in the journey easier to trust.

A strong experiment in any of these sectors follows the same shape. Start with a clear hypothesis, keep one primary metric, run the test long enough to trust the result, and make the decision only after the data has earned it. The sector changes, but the discipline does not.

Tests that connect acquisition to the on-site experience usually beat siloed changes. If the ad promise and the landing page story do not match, the test is fighting a different problem than the media team thinks it is solving.

Pitfalls That Wipe Out Your Test Budget

Most bad tests fail because the setup was bad, not because the idea was weak. The classic mistake is stopping early when the variant looks ahead, then shipping a result that wouldn't survive another week of traffic. Another is changing too many things at once, which makes it impossible to know what caused the shift.

A list of four common pitfalls that can waste an A/B testing budget for digital marketing campaigns.

The operational mistakes that hurt the most

Ignoring seasonality is another easy way to misread the data. A campaign that looks weak during a quiet stretch can look completely different during a promotional period, holiday spike, or weekend-heavy buying cycle. That's why timing matters as much as the page itself.

Low-traffic markets create a different problem. The team may wait too long for significance, get impatient, and then quit before the test has a fair shot. That's not a data problem, it's a planning problem.

  • Stopping Early: Don't freeze a winner before the test has had time to reflect normal traffic patterns.
  • Testing Multiple Changes at Once: Keep the experiment focused so the result tells you something useful.
  • Ignoring Seasonality: Compare like with like, or the result will mislead you.
  • Insufficient Sample Size: If the audience is too small, the test may only produce noise.

Questions to ask before launch

  • What decision will this test inform? If there's no rollout decision tied to the result, the test probably isn't worth running.
  • What's the smallest lift that justifies action? This keeps the team honest about effort versus gain.
  • Can the traffic support the test window? If not, the answer may need to wait.
  • Which segment matters most? Device, source, and region can all change the interpretation.

A good testing calendar is less about being busy and more about avoiding expensive confusion.

The teams that win with experimentation are usually the teams that say no to rushed conclusions. They don't need every test to be dramatic. They need every test to be clean enough to trust, and every conclusion to be tied to a real business choice.

How an Integrated Testing Programme Lifts Revenue

A/B testing has the most value when it's connected, not isolated. The strongest programmes link ad creative, landing pages, forms, and checkout behaviour so the message people click on matches the experience they get after the click. That's how a marketing team moves from “we got more clicks” to “we got better revenue.”

An integrated approach also makes reporting cleaner. If a Meta creative test wins a certain angle, the landing page should reflect that same promise instead of forcing visitors to decode a new message. The same logic applies across Google, TikTok, and LinkedIn, where the ad and the page need to work like one system instead of two separate campaigns.

What mature programmes look for

The best agency setups don't stop at conversion rate. They look at whether the test supports paid acquisition efficiency, because a page that converts better but attracts the wrong lead or buyer can still hurt the business. That's why the conversation should always include the hypothesis, the metric, the sample size, and the rule for shipping.

For a closer view of on-site optimisation thinking, what conversion rate optimisation really means is useful because it frames tests as part of a revenue engine, not isolated page tweaks. In practice, that mindset is what helps teams link a winning ad promise to a landing page that converts on the same promise.

Some published agency case studies point to outcomes like +1250% Meta conversions, +580% revenue growth, and 29% higher conversion rates, which shows the scale an integrated testing model can reach when the acquisition and onsite teams are aligned. The numbers matter, but the operating principle matters more, every test has to fit into a system that keeps learning after the first win.

The right partner should be able to explain that system clearly. If they only talk about creative ideas, they're not running a growth programme. If they only talk about traffic, they're not closing the revenue loop.

Quick Answers to Common A/B Testing Questions

How long should an A/B test run

Long enough to capture normal behaviour, not just a short-lived spike. Nielsen Norman Group recommends at least 1 to 2 weeks as a starting point, and lower-traffic or seasonal markets often need longer if the sample has not stabilised yet.

Can small traffic volumes still produce useful results

Yes, but the question has to be sharper. Small audiences are better suited to bigger, more obvious changes, because tiny lifts are much harder to prove with confidence when the sample is thin.

What if two variants look tied

Treat that as a sign the test did not create enough separation, not as permission to guess. If the difference is not meaningful enough to justify the implementation effort, keep iterating or test a more distinct idea.

How does A/B testing fit with paid media budgets

It helps you avoid paying for traffic that lands on an unproven page. The cleaner the test between ad promise and landing-page experience, the easier it is to protect ROAS while you scale the winners.

If you want a partner that treats testing, media, and conversion rate optimisation as one revenue system, speak to Market With Boost and book a discovery call to review your traffic, identify the biggest leaks, and build a testing plan around the pages that can move revenue.

Justine Bowman

Written by

Justine Bowman

Account Lead

Justine brings over 15 years of agency experience to Boost, with a strong background in traffic management and client operations. She developed her skills at Saatchi & Saatchi BrandsRock, where she learned to keep projects on track, manage client relationships, and deliver campaigns on time.

Hannah Merzbacher photo

Scale your performance with data-driven insights

Ready to apply these insights to your business? Hannah can walk you through how we'd approach your specific situation.

Hannah Merzbacher

Operations Manager

Continue Reading

View all Insights
Attribution Modeling: A 2026 Guide to Revenue Growth
29/07/2026

Attribution Modeling: A 2026 Guide to Revenue Growth

Cut through the jargon of attribution modeling with clear model comparisons and a realistic implementation roadmap for revenue growth....

By Chris EdingtonRead
What Is Conversion Rate Optimisation: Boost Your Revenue
16/07/2026

What Is Conversion Rate Optimisation: Boost Your Revenue

Discover what is conversion rate optimisation (CRO) & why it's vital. Our guide explains the process, key metrics, & turning traffic into revenue....

By Justine BowmanRead
What Is a B2B Ad Agency? a Revenue-Focused Guide
01/07/2026

What Is a B2B Ad Agency? a Revenue-Focused Guide

Looking for a B2B ad agency? This guide explains their services, how to choose one, and why focusing on revenue—not just leads—is the key to growth....

By Justine BowmanRead