The CMO’s Guide to A/B Testing in 2026

How to build an experimentation system when AI can generate 100 ideas before your team can validate one.

12
Chapters
184
Findings
30k+
Tests behind them
01 Intro

The new bottleneck

Marketing teams have a strange new problem.

It has never been easier to make things.

Claude can rewrite your homepage in seconds. AI coding tools can rebuild it in an afternoon. Your team can generate 50 headline variants, three pricing page concepts, a new onboarding flow, and six landing pages before lunch.

The bottleneck has shifted from building to prioritizing what to build.

That changes the role of experimentation.

For years, A/B testing was mostly viewed as an optimization function: take something that exists, make a variation, measure whether conversion goes up.

In 2026, we think that definition is too narrow.

Experimentation increasingly needs to become the validation layer for marketing.

02 The shift

Why this guide exists

  • Competitive advantage is no longer the ability to create and ship more tests (and by the way you would be blown away by how many Fortune 500 website teams have KPIs purely based on number of tests launched)
  • Teams today need to prioritize which ideas will be impactful.
  • Welcome to the Validation Economy.
Read more LLMs give terrible website advice
03 Approach

Taste-driven optimization

  • Most marketing organizations today fall somewhere between three models.
  • This is still how a surprising amount of website work gets done.
  • Analytics show a problem.
  • The team brainstorms solutions.
  • Someone looks at competitors (often grabbing inspiration from a random version of the site that could be a losing version shown to 10% of site traffic)
  • Marketing likes one direction. Design likes another. The founder has opinions.
  • Eventually somebody wins the debate (usually it’s the founder) and something ships.
  • There is nothing inherently wrong with intuition. In a world of AI, taste or intuition is a valuable differentiator.
  • The problem is treating intuition as evidence.
  • We see teams spend weeks debating relatively arbitrary choices while much bigger opportunities sit untouched.
  • And many experiments aren’t actually testing meaningfully different ideas.
  • “Book a Demo” vs. “Get a Demo.”
  • One version of essentially the same headline against another.
  • Or a slightly different testimonial.
  • Then the experiment comes back flat, and everyone concludes:
  • “A/B testing doesn’t work for us.”
  • But often the problem is what tests are actually being run.
04 Approach

AI-accelerated optimization

  • Now we have a newer version of the same problem.
  • Instead of asking the people in the room what they think, we ask an LLM.
  • “Analyze this homepage and tell me how to improve conversion.”
  • And it confidently gives us 15 recommendations.
  • Add social proof.
  • Strengthen the CTA.
  • Make the value proposition clearer.
  • Reduce friction.
  • Personalize the experience.
  • The recommendations sound reasonable because LLMs are very good at generating reasonable-sounding answers.
  • But that’s different from knowing whether those answers work.
  • LLMs largely learn website advice from the information available to them: articles, frameworks, conversations, opinions, and other published material.
  • They generally don’t have access to the private outcomes of the thousands of experiments companies actually run.
  • That’s a critical distinction.
  • An LLM can generate hypotheses. It should not automatically become your source of truth.
  • AI dramatically increases the number of things a marketing team could do.
  • Without better validation, it can also dramatically increase the number of bad things you ship.
05 Approach

Evidence-grounded experimentation

  • This is where we think marketing is headed.
  • You still start with first-party information:
  • → Analytics
  • → Funnel data
  • → Customer research
  • → Heatmaps
  • → Sales conversations
  • → Brand strategy
  • Analytics and research help you see where the problems are but not how to fix them.
  • Before deciding what to build, we recommend adding another layer:
  • What have we already learned from experiments happening elsewhere?
  • Has this idea been tested before?
  • Does it generally create meaningful movement?
  • In what contexts?
  • What tends to win?
  • What tends to lose?
  • What are the exceptions?
  • Instead of starting every experiment from zero, you start with a prior.
  • For example, looking across many experiments can reveal patterns that would be almost impossible to infer from a single test.
  • At DoWhatWorks, we repeatedly see examples where studying many related experiments helps separate a broader pattern from the noise of an individual result.
  • This does not eliminate A/B testing.
  • But the goal moves to, “There is evidence this is worth testing, so let’s spend our limited traffic and engineering resources here rather than on something arbitrary.”
  • That is a fundamentally different experimentation system.
Validate your next idea Pressure test your idea against tens of thousands of A/B tests on your same concept
06 The framework

The evidence test

  • Before your team commits resources to its next website experiment, run the idea through these five tests.
  • Question: Why do we believe this might work?
  • There are many legitimate answers:
  • → Customer interviews
  • → Funnel data
  • → Behavioral analytics
  • → Previous experiments
  • → Competitive experiment data
  • → User research
  • But there should be an answer.
  • Too often, the belief that something will work comes from anecdotal experience or generic LLM recommendations.
  • Those can generate a hypothesis, but they aren’t evidence by themselves.
Watch for A beautifully constructed rationale that ultimately traces back to someone’s opinion.
07 The framework

The meaningful-change test

  • Question: If this variant wins, will we understand why?
  • One of the simplest frameworks we’ve found for thinking about experiments is asking whether the change does at least one of two things:
  • 1. Does it introduce new information?
  • or
  • 2. Does it introduce new direct value?
  • Consider the difference between:
  • “Start Free Trial”
  • vs
  • “Start Your Free Trial”
  • That’s the same exact concept in both versions.
  • But if instead you add reassurance text below your call-to-action button that says…
  • “14-day free trial. No credit card required. Free migration included.”
  • The second version gives the prospect information they didn’t have before.
  • Likewise, personalization that actually changes the experience for someone introduces new value.
  • A great example we point to is MongoDB and Sage, who have toggles where you can select your industry or role and get a customized page based on that.
  • These are materially different experiences.
  • The bigger the conceptual difference between your control and variant, the more likely the experiment is to teach you something meaningful.
The Sage home page with the Explore customized solutions bar ringed and arrowed, offering industry and company size pickers that change the page
Sage lets a visitor pick their industry and company size, which then creates a unique site experience based on selections
Watch for Running dozens of variants that are different in execution but identical in substance.
08 The framework

The risk test

  • Question: What’s the downside if we’re wrong?
  • Not every website decision deserves the same process.
  • Imagine a simple matrix:
  • Low risk / low impact
  • Small visual refinements.
  • Low risk / meaningful upside
  • Replacing vague CTA language with clearer destination-oriented language.
  • High risk / high potential impact
  • Replacing your traditional homepage hero with an AI input interface.
  • Those decisions shouldn’t be treated identically.
  • We have seen changes where the aggregate evidence suggests very little downside and some meaningful upside.
  • Specific CTA language is one example: sometimes the impact is modest, but the implementation cost and downside can also be extremely low (a popular example I share is replacing “Learn More” with something more specific like “Read the Research Paper”. In about 75% of cases it’s neutral, but the rest it’s almost always positive. So very lew risk)
  • Meanwhile, radical changes to user paths deserve considerably more validation.
A two by two chart plotting risk against impact, with changing a colour scheme low on both, replacing a Learn More CTA at low risk and medium impact, and replacing the hero with an LLM box high on both
Low risk and medium impact is where the compounding gains sit, not the high risk corner.
Watch for Building a heavyweight experimentation process that slows down obvious low-risk improvements or casually shipping high-risk changes because they look innovative (the bandwagon effect is real)
09 The framework

The context test

  • Question: How similar is the evidence to our situation?
  • There are very few universal website rules.
  • A tactic that works for:
  • Enterprise SaaS
  • may not work for:
  • Consumer e-commerce.
  • Something that works on:
  • A high-intent pricing page
  • may not work on:
  • An educational blog post.
  • Traffic source matters.
  • Company maturity matters.
  • User intent matters.
  • Page type matters.
  • This is why individual case studies can be misleading.
  • One company’s test is interesting.
  • Dozens or hundreds of related experiments tell you much more.
  • Any “best practices” in conversion rate optimization should really be looked at through the lens of:
  • “This appears to increase our prior probability that this will work under these conditions.”
  • We’ve seen this repeatedly when looking across groups of related tests: patterns emerge, but nuance remains important.
Watch for Turning “best practices” into commandments.
Validate my idea See context before your next site experiment
10 The framework

The learning test

  • Question: What will we know after this experiment that we don’t know today?
  • This may be the most important question.
  • A good experiment doesn’t only produce a winner. The best teams I know running the top websites are always looking to build their foundation of “best practices” that are relevant and replicable to their org.
  • Imagine testing:
  • Headline A vs. Headline B.
  • Headline B drives 7% more clicks.
  • Great.
  • But the takeaway can often be a bit isolated to that specific page.
  • Now compare that with:
  • Generic positioning vs. positioning around a specific use case.
  • If the use-case version wins, you have potentially learned something applicable to:
  • → Your homepage
  • → Paid acquisition
  • → Sales collateral
  • → Onboarding
  • → Email
  • → Product positioning
  • The value of that experiment extends far beyond its immediate conversion lift.
Watch for Optimizing individual pages without building an institutional understanding of why customers respond.
11 Convictions

What we believe

  • After looking at experiments across many of the world’s largest companies, a few beliefs have become central to how we think about testing.
  • 1. The explosion in creation makes validation more important
  • AI is not reducing the need for experimentation.
  • It is increasing it.
  • When generating another page costs almost nothing, organizations can easily create more ideas than they can intelligently evaluate.
  • The scarce resource becomes confidence.
  • We think validation will increasingly become a core layer between idea and deployment.
  • 2. LLMs are incredible generators and imperfect judges
  • We use AI. We love AI. It’s an insane accelerant.
  • But generation and validation are different jobs.
  • Use AI to:
  • → Generate hypotheses
  • → Explore alternatives
  • → Analyze research
  • → Draft variants
  • → Prototype experiences
  • Then ask:
  • What evidence suggests any of these ideas will actually improve the outcome we care about?
  • An articulate answer is not necessarily an accurate answer.
  • 3. Stop spending your testing budget on insignificant questions
  • Traffic is finite.
  • Engineering time is finite.
  • Design resources are finite.
  • Organizational attention is finite.
  • Every trivial test has an opportunity cost.
  • The best experimentation teams are running tests on the most impactful aspects of their digital experience.
  • 4. Patterns matter more than anecdotes
  • Our CMO, Casey Hill, frequently highlights creative website ideas that sites launch.
  • It makes for great Linkedin content. It can help provide inspiration.
  • But learnings are anecdotal on an individual level.
  • Fifty related experiments, though, start to tell you something about human behavior.
  • This is especially important because individual tests contain enormous amounts of noise: implementation differences, audiences, traffic sources, timing, and countless other variables.
  • Our conviction increases when we repeatedly observe the same behavior across companies and contexts.
  • And the more often those tests are isolated to a single variable being tested (A/B testing 101 here), the more explanatory value it has.
12 Where this lands

The bottom line

  • For the last decade, marketers have been told to move faster.
  • AI solved that problem.
  • We can now produce ideas, campaigns, copy, and digital experiences at a speed that would have been unimaginable a few years ago.
  • But moving faster in the wrong direction isn’t progress.
  • The companies that win the next era of marketing won’t just be the companies that can build the fastest.
  • They’ll be the companies that can learn the fastest.
  • And learning requires something AI alone cannot give you:
  • Evidence.
  • That’s why we believe the next era of experimentation is not about running more A/B tests.
  • It’s about building a better system for determining what works.
Join Alpha Build a validation layer into your testing program