Every performance marketer has had the same uncomfortable thought while looking at a dashboard. The numbers look good. Reported ROAS is healthy. Conversions are up. And somewhere at the back of your mind sits the question you cannot answer from that screen: how many of these sales would have happened anyway?
Incrementality testing is how you answer it. It is the practice of measuring the conversions your advertising actually caused, rather than the conversions your advertising happened to be present for. In 2026 this has stopped being an advanced technique and started being basic hygiene, because automated bidding, modelled conversions and privacy-driven data loss have all made platform reporting more confident and less verifiable at the same time.
What incrementality actually means
Incrementality is the share of your results that would not have happened without the advertising. Everything else is a conversion you paid to observe rather than a conversion you paid to create.
The mechanics are simple. You take two comparable groups. One sees your ads, one does not. You let them run for a set period. Then you compare outcomes. The difference between the two groups is your incremental result, and dividing your spend by that difference gives you a cost per incremental conversion that is grounded in something other than platform self-reporting.
The concept is borrowed from clinical trials, and the comparison is useful. A drug trial does not ask patients whether they felt the medicine helped. It gives one group the treatment, gives another a placebo, and measures the gap. Ad platforms, by contrast, are in the awkward position of being both the treatment and the lab reporting the results.
A quick example
A retargeting campaign reports 500 conversions at a 6x ROAS. You pause it in half your markets for four weeks. Sales in the paused markets drop by 90 conversions relative to the markets that kept running. Your incremental contribution is 90, not 500. The other 410 people were coming back regardless. Your true cost per acquisition on that campaign is roughly five and a half times what the dashboard told you.
Incrementality vs attribution: two different questions
Attribution and incrementality get treated as rivals. They are not. They answer different questions, and confusing the two is the single most common measurement mistake we see in paid media teams.
Attribution is a bookkeeping exercise. It observes conversions that already happened and distributes credit across the touchpoints it can see. Whether you use last click, data-driven, or something you built yourself, the model is describing correlation. If you are still deciding which model to run, our guide to marketing attribution models covers the trade-offs. Incrementality is an experiment. It creates a counterfactual on purpose so you can measure causation.
| Dimension | Attribution | Incrementality |
|---|---|---|
| Question answered | Which touchpoints were involved? | What did the spend cause? |
| Type of evidence | Correlational | Causal |
| Availability | Continuous, in every dashboard | Periodic, requires a deliberate test |
| Granularity | Down to keyword and creative | Usually channel or campaign level |
| Typical bias | Overstates lower-funnel channels | Understates long-cycle effects |
| Best used for | Day-to-day optimisation | Budget allocation decisions |
The practical division of labour: use attribution to steer inside a channel, use incrementality to decide how much that channel deserves. Optimising creative rotation with a geo holdout would be absurd. So is deciding next year's budget split on last-click data.
Four incrementality test types worth your time
There are many ways to build a counterfactual. Four of them are practical for a team without a dedicated measurement function.
1. Geo holdout tests
Split your market into comparable regions. Keep advertising in one set, switch it off in the other, then compare total sales rather than platform-reported sales. This is the most robust option because it measures the whole business outcome and it survives cookie loss, ad blockers and cross-device journeys. It needs geographic spread and enough volume per region. Meta's open source GeoLift library and Google's CausalImpact both handle the statistics if you can work in R.
2. Platform conversion lift studies
Meta, Google and the larger platforms will hold back a randomised share of your audience and report the lift for you. The split is clean and the setup takes minutes. The obvious caveat is that the platform grades its own homework, and the measurement stops at the edge of that platform. Useful, but not the only evidence you should accept.
3. Budget-split experiments
Rather than switching a campaign off, run it at two different spend levels across randomly assigned traffic. Google Ads campaign experiments do this natively. You learn the shape of the curve, which is the question that usually matters most: not whether the channel works, but whether the next euro works.
4. Time-based on and off tests
Alternate weeks or fortnights with the channel on and off, then compare periods. The weakest of the four, because seasonality, competitor activity and promotions all contaminate the comparison. Use it only when geography and audience splits are impossible, and interpret the result with appropriate caution.
Which one should you pick?
- arrow_forwardMultiple regions and healthy volume. Geo holdout. It gives you the cleanest read on real business impact.
- arrow_forwardSingle market, one dominant platform. Platform conversion lift, treated as directional rather than definitive.
- arrow_forwardYou already know the channel works. Budget-split experiment, to find where the returns start flattening.
- arrow_forwardLow volume, long sales cycle. Test at channel level over a longer window, and accept a wider confidence interval.
How to run an incrementality test in five steps
Most failed tests fail in the design phase, not the analysis. Spend the time up front.
Write down the decision the test will inform
Not the hypothesis, the decision. "If branded search proves less than 30 percent incremental, we move half that budget to prospecting." A test without a pre-agreed consequence produces an interesting slide and no change in behaviour.
Pick one metric and one variable
One primary outcome, ideally revenue or qualified leads rather than a proxy. One thing changing between groups. If you also refresh creative or shift bids mid-test, you have destroyed your own control.
Size the test before you launch it
Work out how many conversions you need to detect the effect you care about. Detecting a 20 percent lift takes far less volume than detecting a 5 percent lift. If the maths says your account cannot produce a readable answer in eight weeks, test at a higher level or do not run it.
Establish a clean baseline and then leave it alone
Capture four to eight weeks of pre-period data so you can verify the groups behaved similarly before the test. Then hold your nerve. Every bid change, budget nudge and campaign restructure during the window is contamination, and the temptation to intervene peaks exactly when early numbers look bad.
Read the result honestly, including the range
Report the confidence interval alongside the point estimate. "22 percent incremental, somewhere between 9 and 35" is a real finding. "22 percent incremental" on its own invites a precision the data does not support.
On being honest about the limits
An incrementality test is a snapshot of one channel, one period and one set of market conditions. It is not a permanent verdict. Competitor entries, seasonality and creative fatigue all move the number. Treat results as having a shelf life of a couple of quarters, and re-test the channels carrying the most budget.
Reading the results: incremental ROAS and what to do next
Two numbers come out of a well-run test. The incrementality rate, which is incremental conversions divided by reported conversions. And incremental ROAS, which is revenue caused by the campaign divided by the spend.
The second number is the one that changes budget decisions, and it is almost always lower than the figure in your platform reports. If you want a refresher on the underlying calculation, we broke it down in our guide to ROAS calculation. The gap between reported and incremental ROAS is not a reporting error to be fixed. It is a measurement of how much of your budget is buying demand that already existed.
What the numbers usually tell you to do
- arrow_forwardHigh incrementality, strong incremental ROAS. Scale it, and run a budget-split test to find where the curve flattens.
- arrow_forwardLow incrementality, strong reported ROAS. The classic retargeting and branded search pattern. Reduce spend in steps and watch total revenue rather than campaign revenue.
- arrow_forwardHigh incrementality, weak reported ROAS. An underrated upper-funnel channel that attribution is starving. This is where teams find their biggest wins.
- arrow_forwardResult within the noise. Not a failure. It usually means the effect is smaller than your test could detect, which is itself an argument against the current budget.
One warning before you act. Do not feed incremental conversions back into smart bidding as your optimisation signal. The algorithms need volume, and starving them of conversion data to make the reporting more honest will make the bidding worse. Keep the two separate: platform signals for bidding strategies, incrementality for allocation.
Five mistakes that make incrementality tests useless
Stopping early
Checking results daily and calling it when the gap looks convincing is how you generate false positives at scale. Set the end date before you start and honour it.
Contaminated control groups
Your held-out region still sees your email, your organic listings and your YouTube. Some spillover is unavoidable. Document it so the result gets read as conservative rather than clean.
Measuring platform conversions
If your outcome metric comes from the same platform you are testing, you have not escaped the problem. Use your own backend revenue data as the source of truth.
Testing everything at once
Running four tests across overlapping audiences in the same quarter means none of them has a clean control. Queue them instead, and pick the channel with the most budget at stake first.
- infoAnd the fifth: shelving the result. The most common failure is not statistical at all. A team runs a good test, learns that a large slice of spend is not incremental, and then quietly does nothing because the reallocation is politically awkward. Agree the decision rule before the data arrives and this problem mostly disappears.
Where to start this quarter
If you have never run one, do not begin with a complicated multi-cell design. Pick the campaign with the highest reported ROAS and the loudest internal reputation. In most accounts that is branded search or retargeting, and in most accounts it is also where the reported number is furthest from the real one.
Run a single geo holdout for four weeks. Measure total revenue, not campaign revenue. Whatever the result, you will have learned something no dashboard could have told you, and you will have a repeatable process for the next channel.
The prerequisite for all of this is boring: knowing what you are spending, where, and whether it is pacing to plan. Testing is impossible when the underlying spend data lives across six tabs and a spreadsheet someone updates on Fridays. If that is your situation, fix the tracking before you fix the measurement, starting with a reliable way to track ad spend across channels.
Frequently Asked Questions
Know what your budget is doing, every day
aubado keeps spend, pacing and performance across every channel in one place, so the only thing left to decide is what to test next. Check once a day. Then close the tab.
Share this article