Running an Incrementality Test When You Cannot Afford to Switch Spend Off
How to get a real read on incremental sales when your budget is too small for a clean holdout, and which designs are honest at that size.
- Author
- Prabhash Jha
- Published
- Reading time
- 16 min read
Every guide to incrementality testing gives the same instruction, and it is correct: pause the channel in a treatment region, leave it running in a comparable control region, and read the difference. Do not lower budgets — actually pause them, because a reduced budget still serves impressions and contaminates the holdout. Four weeks minimum. Then it tells you the minimum spend that makes the test viable, and for a large number of businesses that number is above what they spend in total.
Which leaves the honest version of the question, the one nobody writing measurement software has any commercial reason to answer: what do you do when the correct test is out of reach? You still have to decide whether the channel is worth what you are paying for it. The last-click number in the platform is not going to tell you, because that number was never designed to.
The answer is not a smaller version of the proper test. A geo holdout run on a budget that cannot support it does not produce a slightly noisier result — it produces a confident number that is mostly noise, which is strictly worse than no number, because you will act on it. The answer is a different set of designs that trade statistical strength for something you can actually run, plus a much clearer idea of which decisions the weaker evidence is allowed to make.
Why does the platform’s ROAS overstate incrementality?
Because the platform counts conversions it influenced and conversions it merely observed, and it cannot tell them apart.
This is not a criticism of the platform so much as a description of what it is measuring. An ad platform reports conversions where its ad was in the path. Some of those buyers were going to buy anyway — they were already searching for you, already had the app installed, already had it in the cart. The ad appeared, they clicked it because it was there, and the sale got attributed. Incrementality is the difference between what happened and what would have happened with the spend switched off, and no attribution model can see the counterfactual, because the counterfactual did not happen.
The size of the gap is the surprising part, and the best evidence on it comes from someone who could actually run the experiment at scale. Blake, Nosko and Tadelis ran a series of large field experiments at eBay, published as Consumer Heterogeneity and Paid Search Effectiveness (NBER working paper, later in Econometrica). Their headline finding is that returns from paid search were a fraction of the conventional non-experimental estimates — and in the extreme case, brand-keyword ads had no measurable short-term benefit at all. For non-brand keywords they found new and infrequent users were positively influenced, but frequent users, whose behaviour the ads did not change, accounted for most of the spend, producing negative average returns.
Two things are worth taking from that and one thing is worth being careful about. Take: the direction of the bias is always the same, and the mechanism is substitution — the paid click replaces a click that was free. Take: the effect is heavily concentrated in your most loyal, most frequent buyers, which is precisely the audience that looks best in the platform’s report. Be careful about: eBay is one business with an enormous organic brand presence, and the magnitude does not transfer. A new brand nobody searches for by name has a very different substitution profile, because there is far less free traffic for the paid click to substitute for. The lesson to carry across is that the gap exists and is large enough to change decisions — not that your own number is inflated by the same amount as theirs.
If you have not yet separated the two ideas cleanly, attribution windows are a commercial term, not a technical setting is the companion to this — that post is about who gets credit, this one is about whether the credit was earned at all.
What is the actual minimum for a valid geo holdout?
There is no fixed rupee figure, and anyone quoting one is guessing. The binding constraint is conversions per region per week, not spend.
This is the most useful correction to make early, because it changes who can run the test. A geo holdout compares two groups of regions and asks whether the treated group’s outcome diverged from the control group’s. The precision of that comparison is driven by how many conversions each group produces and how volatile that number normally is. A business with a high average order value and eight sales a week cannot run this test at any budget. A business with a low order value and hundreds of transactions a week may be able to run it on a small one.
So the pre-test question is not “can I afford this” but:
| Check | Why it decides the test |
|---|---|
| Conversions per week, per candidate region | The unit of statistical power. Few conversions, no test |
| Week-to-week volatility of that number | Your effect has to be bigger than the normal wobble |
| How much of the effect you expect | A channel you suspect is 90% non-incremental is easy to detect. One you suspect is 15% off is not |
| Whether regions behave alike historically | If control and treatment already drift apart, the test measures the drift |
| Whether the channel can even target geographically | Some cannot, at which point the design is unavailable regardless |
The methodology behind all of this is public and worth reading once rather than taking on trust from a vendor: Vaver and Koehler’s Measuring Ad Effectiveness Using Geo Experiments is the paper most geo-testing products are built on top of, and it is clear about the design process and about what makes a result interpretable. Reading it also inoculates you against the version of the pitch that treats “we ran a geo test” as self-evidently rigorous.
The fourth row is the one that quietly fails most small tests. Two regions that look similar on a map are not similar for this purpose unless their historical conversion series moved together. Before you run anything, plot the last six months of weekly conversions for both candidate groups on the same axes. If they already diverge for reasons you cannot explain, the experiment is going to hand you that divergence and call it your result.
Designs that work when a clean holdout does not
Below are the four I would consider, in descending order of how much they are worth. Each one is weaker than the proper test. Each one is dramatically better than reading the platform report and calling it measurement.
1. Reverse the holdout: hold out a small region, not a large one
Instead of pausing a large share of your spend, pause a small, well-behaved region and read the loss there.
The instinct with a limited budget is that a holdout means giving up a big chunk of revenue for a month. It does not have to. If you pause one region that represents a modest slice of volume, you are risking that slice and nothing else. The trade is that a small region produces fewer conversions and therefore a noisier read, so this only works if that one region still clears the conversion threshold above. When it does, it is the highest-quality evidence available to a small advertiser, because it is a genuine experiment rather than an inference.
The variant most people miss: you are allowed to run it long instead of big. Statistical power comes from total conversions observed, and a smaller region observed for eight weeks can carry as much information as a bigger one observed for three. What you pay for that is calendar time and the risk that something else changes mid-test — which is a real cost, not a free lunch, and it is why this design suits stable months and not the run-up to a sale.
2. Read what happens when spend stops for a reason that is not your test
Use the outages, budget exhaustions, card failures and account suspensions you already have.
Every account has these. A card declines and the campaign stops for two days. A budget caps out mid-month. An account gets suspended over a policy warning and comes back on Thursday. Each of those is an unplanned, involuntary interruption to spend, and each one is a small natural experiment you did not have to pay for. Nobody logs them, so nobody can read them.
Start the log now. A spreadsheet is fine: date, channel, what stopped, how long, and total orders on those days against the same weekday in the surrounding weeks. Individually each event is nearly worthless — two days of data, confounded by whatever else happened that week. Collected over a year, the pattern is genuinely informative, and it costs you nothing but the discipline of writing it down. When a channel goes dark four separate times and total orders do not visibly move on any of them, that is evidence, and it arrived free.
The caution that goes with it: these interruptions are not random. Cards fail at month end, budgets exhaust in high-demand periods, suspensions follow aggressive creative. Each of those correlates with something. So this is a signal to weigh, not a result to quote, and if it disagrees with a real experiment the experiment wins.
3. Move the spend instead of removing it
Reallocate the same money between regions rather than switching it off, and read the difference in both directions.
If your total budget must stay constant — which is the usual real constraint, because the budget is a commitment to a client or a board rather than a dial — you can still create variation. Increase spend meaningfully in region A, decrease it by the same amount in region B, hold everything else. You now have two effects to observe rather than one, in opposite directions, funded by each other.
This is a genuinely weaker design than a holdout, for a reason worth understanding: it measures the effect at the margin you moved, not the total contribution of the channel. Learning that the next rupee does little is not the same as learning that the whole channel does little. Response curves flatten. The last 20% of a budget frequently does almost nothing while the first 20% does a great deal, and a marginal test cannot distinguish “this channel is worthless” from “this channel is saturated”. For a budget decision — should I spend more here — that distinction does not matter and this test answers it directly. For an existence decision — should this channel exist — it does matter, and this test cannot answer it.
4. Ask the buyer, and treat the answer as weak but real
Add a “how did you hear about us” field at checkout, expect it to be badly wrong, and use it anyway.
Self-reported attribution is genuinely unreliable. People misremember, they name the last thing they saw, they pick the first option in the list, and a meaningful share skip it. Nobody should run a budget off it. But it has one property no platform report has: it is not generated by the platform that gets paid when it looks good. It is independent, and independent-and-noisy is a useful complement to precise-and-biased.
Use it for divergence, not for levels. If a channel reports 30% of your conversions and essentially no customer ever mentions it unprompted, that mismatch is worth investigating. It is not proof. It is a place to point the next real test.
The affiliate case, which is the sharpest version of this
For affiliate and partner channels, the incrementality question is not academic — it is the difference between a fair commission and paying for your own customers.
Everything above applies, but affiliate has a specific and much more tractable structure, because the suspicion is usually about a particular partner rather than the whole channel. The partners that raise it are recognisable:
- Coupon and cashback sites. The user is at your checkout, opens a new tab, searches ”
coupon”, clicks, and returns with a cookie set. The partner has intercepted a sale that was already happening. Some coupon partners genuinely recruit new buyers; the ones sitting purely on your brand terms mostly do not. - Partners bidding on your brand name. The paid-search substitution effect from the eBay study, except you are paying commission on top of it. This has a whole post of its own — your top affiliate is bidding on your own brand name — because the fix is contractual rather than statistical.
- Toolbar, extension and retargeting partners whose entire model is late-funnel presence.
The good news is that the test is much easier here than for paid media, because you can pause one partner rather than one channel, and one partner is a small enough share of revenue that the downside is affordable. Pause a single suspected partner for four weeks and watch total orders, not the partner’s own reported conversions. Three outcomes:
| What total orders do | What it means | What to do |
|---|---|---|
| Fall by roughly what the partner was reporting | The partner is incremental | Keep the rate, possibly raise it |
| Do not move at all | The partner was intercepting | Renegotiate the rate or the placement rules |
| Fall by some fraction | The usual answer | Reprice towards the fraction, not the report |
That middle row is a commercial conversation and not necessarily a termination — a partner who converts sales you would have got anyway is still doing something, and the right response is often a lower rate for that traffic type rather than removal. What it is not is a partner worth the same commission as one bringing you a buyer who had never heard of you.
Before you run any of this, confirm your tracking is actually sound, because an incrementality result computed on broken tracking is a confident answer to the wrong question. Affiliate tracking breaks quietly covers what to check and when.
Which decisions is a weak test allowed to make?
This is the part that matters more than the methodology, and it is almost never stated: match the strength of the decision to the strength of the evidence.
The failure mode with limited-budget testing is not running a weak test. It is running a weak test and then treating its output as though it came from a strong one, because by the time you have a number the caveats have fallen off it. So decide the rule before you see the result:
| Evidence you have | Decisions it can support | Decisions it cannot |
|---|---|---|
| Platform report only | Creative and audience iteration within a channel | Anything comparing channels or setting total budget |
| Self-reported survey + natural outages | Where to point the next real test | Killing a channel |
| Marginal reallocation test | Should the next increment go here or there | Whether the channel should exist |
| Small-region holdout, one run | A meaningful budget shift, reversible | A permanent restructuring |
| Repeated holdouts agreeing | Structural change: kill it, or double it | — |
The bottom row is the point of all this. A single small test is weak. A single small test repeated three times over six months, agreeing each time, is not weak at all, and it is achievable on a budget that could never fund one big clean experiment. Incrementality on a small budget is a programme, not a project. The businesses that get real answers are not the ones that found a clever design; they are the ones that kept a log and ran the same modest test four times.
The mistakes that produce a confident wrong answer
Reducing budget instead of pausing. A reduced budget still serves impressions to your highest-intent users, which is exactly the population whose behaviour you are trying to observe. The contamination is not proportional to the reduction — it is concentrated where the effect lives.
Testing during a period that is not normal. A festive season, a sale, a competitor’s launch, a PR moment. The test will faithfully measure that event.
Stopping early because the answer looks clear. Conversion series are volatile and the first ten days will look decisive in one direction or the other roughly whichever way the channel actually works. Set the duration in advance and do not read it daily. If you must look, agree beforehand what would make you stop, and that it must be a problem, not a result.
Reading the channel’s own reported conversions during the test. During a partner or channel pause, the only number that means anything is total orders. The paused thing will obviously report fewer conversions. That is not the finding.
Forgetting the lag. If your purchase cycle is three weeks, the first three weeks of a pause still contain sales driven by the earlier exposure, and the first three weeks after it restarts still look flat. Both ends of the test are contaminated by the consideration window. Read the middle, and make the test longer than the cycle.
Running it on tracking you have not checked. Same class of error as your Search Console numbers are lying to you — the measurement instrument has its own behaviour, and assuming it is a clean window onto reality is how a data problem becomes a strategy problem.
FAQs
Can I use the platform’s own lift test instead? Yes, and you should, with one reservation. Platform-run lift tests are proper randomised experiments and far better than attribution reports. The reservation is structural rather than an accusation: the platform designs the test, defines the conversion, sets the window and reports the result, and it is paid on the outcome. Use it, and treat a result that disagrees with your own holdout as a question rather than a verdict.
How long does a test need to run? Longer than your consideration window, and long enough to accumulate the conversions the design needs — four weeks is the usual floor for lower-funnel campaigns and it is a floor, not a target. If your purchase cycle is long, the honest answer may be that a four-week test cannot see your effect at all.
What if the result says the channel is not incremental? Do not switch it off on one result. Re-run it. Genuinely non-incremental spend is common enough that the finding is plausible, but the cost of being wrong is asymmetric — turning off a channel that was working is slow to detect and slower to rebuild, because you lose the learning phase, the audience data and the account history along with the spend.
Does incrementality matter for brand campaigns too? It matters more, and it is much harder to measure, because the effect is slower than any four-week test can see. This is one of the arguments in when performance marketing stops working and brand is the only lever left: a measurement approach tuned to a four-week window will systematically under-credit anything that pays back over a year, and then the budget follows the measurement.
Where does this fit against ordinary campaign reporting? Reporting tells you what happened inside a channel and is the right tool for creative and audience decisions. Incrementality tells you whether the channel is producing sales that would not otherwise exist, and is the right tool for budget decisions between channels. Both are needed; the mistake is using the first for the second. The routine version of the first is in the performance marketing playbook I actually use.
Where to start on Monday
If the budget cannot support a proper test, do these three things in this order, because each is cheap and the first two cost nothing at all.
- Start the outage log today. Every unplanned stop, dated, with orders on those days. In twelve months it is the most valuable measurement asset you own, and you cannot backfill it.
- Plot your candidate regions against each other for the last six months. You will learn within an hour whether a geo test is available to you at all, which is worth knowing before you plan one.
- Pause your single most suspicious affiliate or partner for four weeks and watch total orders. This is the highest ratio of insight to risk anywhere in the list, and for most businesses running partner programmes it is the test that changes a number.
None of that is as good as a properly powered geo experiment. All of it is better than the number in the dashboard, which is the actual comparison — not against the ideal test you cannot run, but against the belief you currently hold, which came from a report written by the party being paid.
If you are spending real money on paid and the numbers are not behaving, that is the work I do. See how I work with people.