Ali Fakhar
Measurement

Marketing experiments: use AI to test better hypotheses

Design a marketing experiment that can change a decision: an observation, a falsifiable hypothesis, a fair comparison, guardrails and a volume check before launch. Includes what to do when B2B traffic is too low.

marketing-experiment-design-ai

The short version

  • Start from an observed problem and more than one explanation, not from a list of variations.
  • Define the comparison, unit of assignment, one primary measure and guardrails before seeing data.
  • Check the required sample first. On low traffic, small differences can take years to detect; change the question instead.
  • Use AI to criticise the brief, not to predict the winner or stand in for real results.
On this page

“Let’s A/B test a few headlines” is how a lot of marketing experiments begin, and it is why they often teach nothing. A week later the dashboard shows version B ahead, someone declares a winner, and nobody can say what was learned or whether the result would hold next month.

A useful experiment starts from an observed problem and a clear idea of why it happens. It defines one change, a fair comparison, a primary measure and the guardrails that must not get worse. Before it starts, it checks whether the available traffic can answer the question at all, and it decides what a positive, negative and inconclusive result will each mean. AI is a good critic of the experiment brief. It is not a substitute for the result.

Start from an observation, not a variation

Good experiments begin with something you have noticed in the data or in conversations: visitors reach the pricing page and leave, a campaign gets clicks but no qualified enquiries, prospects on calls misunderstand what the product does. That observation is the reason the experiment exists.

Then ask why it might be happening, and write down more than one explanation. Visitors may leave the pricing page because the price is high, because they cannot tell which plan suits them, or because they came from an ad that promised something else. Each explanation points to a different change. Testing a new headline when the real issue is plan confusion wastes the test.

Write a hypothesis that could be wrong

A hypothesis is a specific, falsifiable prediction with a reason attached. This template keeps it honest:

For [audience], if we [change], then [primary outcome] will [increase or decrease], because [mechanism]. We will judge it by [measure] over [period].

Compare two versions:

  • Weak: “A new headline will improve conversions.”
  • Useful: “For visitors from our operations-software ads, if the landing page headline names the weekly report problem instead of the product category, then the form completion rate will increase, because the ad and page will describe the same problem. We will judge it by completed enquiry forms per unique visitor over the test period, with sales-accepted enquiries as a guardrail.”

The second can be wrong in a way you would notice. That is the point.

Define the comparison, unit, measure and guardrails

Four decisions make or break the design:

  • Comparison. What does the change compete against? Usually the current version, run at the same time, so that seasonality and campaign changes affect both equally.
  • Unit of assignment. Who or what gets randomly assigned: a visitor, an email recipient, an account, a region, a week. In B2B, several people from one account may visit; if they see different versions, the comparison blurs. Sometimes the account is the right unit.
  • Primary measure. One number, defined precisely, that decides the result. Choose it before you see any data.
  • Guardrails. Measures that must not get worse: lead quality, sales acceptance, unsubscribe rate, cost per qualified opportunity. A form that converts better by attracting people sales cannot help is not a win.

Check whether your volume can answer the question

This is the step low-volume B2B teams skip, and the one that matters most for them. Before you run a test, estimate how much data you need to detect the difference you care about.

An illustration, with invented numbers. A landing page receives 1,200 visits a month and 2% of visitors complete the form. The team hopes a new message will raise that to 2.5%. Using a standard sample-size calculation for comparing two proportions, with the common conventions of 5% significance (two-sided) and 80% power, detecting that difference needs roughly 13,800 visitors per version: about 27,600 in total. At 1,200 visits a month, that is almost two years.

Even a much larger improvement, from 2% to 3%, needs about 3,800 visitors per version, over six months of traffic. These figures depend entirely on the assumptions; use a proper calculator with your own baseline and the smallest effect worth acting on. The general lesson holds: small differences on small traffic cannot be detected reliably in a normal campaign cycle.

When the numbers do not work, change the question rather than running the test anyway:

  • Test a bigger change. A different offer or audience is more likely to produce a detectable difference than a new adjective.
  • Pool traffic. Run the change across several pages or campaigns that share the same problem.
  • Use discovery methods instead. Interviews, comprehension tests and sales-call reviews can tell you whether a message is understood. They are valuable, but they are not causal evidence; label them honestly.
  • Accept a directional read. A before-and-after comparison can inform a decision if you record what else changed at the same time and do not claim more certainty than it gives.

Check the data before the result

Before interpreting any result, check that the experiment ran as designed. One simple check: did each version receive the share of users it was supposed to? Microsoft’s experimentation team describes this problem, sample ratio mismatch, as a statistically significant difference between the observed and configured split. They point out that the users who go missing are often the ones most affected by the change, which is why a mismatch can invalidate the result. Causes range from redirects and faulty assignment to data joins that drop records.

If a 50/50 test shows a clearly uneven split, find the cause before reading the outcome.

Use AI as a critic of the brief

AI is useful before the test, when the brief is still cheap to change. Give it your observation, hypothesis and design and ask it to find the weaknesses:

Keep two boundaries. Statistical calculations and assumptions should be checked with an appropriate calculator or a person who knows the method; do not rely on a model’s arithmetic. And a model’s prediction of which version will win, or its simulated customer reactions, is not an experimental result. For generating the creative variations themselves, see how to test ad creative with AI.

Interpret positive, negative and inconclusive results

Decide what each outcome means before you start:

  • Positive: the primary measure improved beyond the threshold and the guardrails held. Roll out, and keep monitoring the guardrails after launch.
  • Negative: the change made things worse. That is useful: the explanation behind it is probably wrong, and you have saved yourself from rolling it out.
  • Inconclusive: no reliable difference. On low traffic, this is a likely outcome. It means the test could not separate the versions, not that they are equal. Decide in advance what you will do: keep the simpler version, run a bolder test, or pursue a different explanation.

Do not stop the test the moment the dashboard looks promising. Repeatedly checking and stopping at the first good-looking moment makes a false “winner” far more likely. Agree the stopping rule before launch.

A completed experiment card

This card is illustrative. The page, audience and numbers are invented.

Completed experiment card (illustrative)
FieldEntry
ObservationVisitors from operations-software ads leave the landing page without enquiring; sales calls suggest the page describes a category, not their problem
HypothesisNaming the weekly-report problem in the headline will raise form completion, because ad and page will describe the same problem
InterventionNew headline and first paragraph; everything else unchanged
ComparisonCurrent page, running at the same time
Audience and assignmentVisitors from the three operations-software campaigns, randomised by visitor
Primary measureCompleted enquiry forms per unique visitor
GuardrailsShare of enquiries accepted by sales; cost per qualified opportunity
Volume checkAt 1,200 visits a month, only a very large lift is detectable; traffic pooled from three campaigns and a bolder message chosen
Stop rulesFixed end date agreed in advance; stop early only for a guardrail breach or broken tracking
Decision if inconclusiveKeep the clearer message on qualitative evidence from calls, and label the decision as not experimentally proven

Repair a flawed brief

Here is a brief with typical problems. Try to spot them before reading the list.

We will test three new headlines against the current one for two weeks and pick whichever gets the most clicks. If it goes well, we will roll it out everywhere.

The problems: no observation or hypothesis, so nothing is learned whatever happens; four versions split already small traffic four ways; “most clicks” is not the business outcome and has no guardrail; the two-week duration is not based on a volume check; “if it goes well” is undefined; and “roll out everywhere” assumes other pages share the same problem.

Blank experiment card

Blank experiment card
FieldEntry
Observation[What you noticed, and where]
HypothesisFor [audience], if we [change], then [outcome] will [direction], because [mechanism]
Intervention[Exactly what changes]
Comparison[What it runs against]
Audience and assignment[Who is included; unit of randomisation]
Primary measure[One precisely defined number]
Guardrails[What must not get worse]
Volume check[Baseline, smallest useful effect, required sample, expected duration]
Stop rules[When the test ends; what stops it early]
Decision for each outcome[Positive, negative, inconclusive]

Experiments sit inside a wider measurement question: does the thing you are testing connect to pipeline at all? That is the subject of B2B marketing metrics: connect lead quality to pipeline. And a message worth testing starts with clear positioning.

Sources and further reading

  1. Microsoft Research: Diagnosing sample ratio mismatch in A/B testing

Questions worth asking

Is every marketing experiment an A/B test?

No. A randomised A/B test is one design. Pooled tests across several pages, before-and-after comparisons with recorded context, and qualitative methods such as comprehension tests are all useful. Only randomised comparisons give strong causal evidence, so label the others honestly.

What if traffic is too low?

Estimate the required sample before starting. If it would take months or years, test a bigger change, pool traffic across pages that share the problem, or use interviews and sales-call reviews for a directional answer. Do not run an underpowered test and treat the result as proof.

Can AI predict the winning ad?

Not reliably, and its prediction is not evidence. A model can generate variations and point out weaknesses in a test design. Whether a message works for your buyers is something only a real test or real buyer feedback can show.

Ali Fakhar
About the author

Ali Fakhar

Ali Fakhar is a London-based marketer working across growth, paid media, content and practical AI.

More about Ali ↗
Analytics preferences

Current preference: off