
Here is the uncomfortable structure of modern paid media. Meta reports on Meta. Google reports on Google. Both are commercially motivated to show you a large number, both count conversions they merely touched rather than caused, and both now hide the campaign level detail behind automated buying systems you cannot fully inspect. You are being asked to allocate six or seven figures a year on the basis of a self-graded exam.
Incrementality testing is the correction. It answers one question that no attribution dashboard can: if this spend disappeared tomorrow, how much revenue would disappear with it? The mechanics are not complicated. The design discipline is where almost every brand falls over, which is why so many operators run a test, get a number they dislike, and quietly go back to trusting the platform.
What incrementality actually measures
Platform attribution counts conversions that can be linked back to an ad interaction inside a lookback window. That is a very different question from whether the ad changed the outcome. A customer who was already going to buy, who saw a retargeting ad on the way to checkout, shows up as a conversion Meta claims. She would have converted anyway. The ad cost you money and bought you nothing.
Incremental ROAS, usually written iROAS, is the ratio of incremental revenue to spend. It is almost always lower than the number on your dashboard, and the gap is not uniform across channels. Retargeting and branded search tend to show the largest gaps because both largely harvest demand that already exists. One vendor reporting across 225 DTC geo tests put the median iROAS on branded search at 0.70x, meaning the average brand in that sample was paying to buy clicks it would have received organically.
Attribution tells you where a conversion has been. Incrementality tells you whether you needed to pay for it. Those are different businesses.
The three ways to test, ranked honestly
Not every test is worth running. These are the three designs you will be offered, from weakest to strongest.
Time based on off test
Run ads, pause ads, compare the two periods. This is the most commonly run design and the least valid one. Between your two periods you will have had a promotion end, a competitor launch, a payday, a weather shift, a press mention, or simply a different point in the month. There is no control group to absorb any of it, so every one of those changes lands inside your result. If someone offers you this as an incrementality test, they are offering you a story, not a measurement.
Platform conversion lift
Meta and Google will both hold out a randomised share of your audience and report the lift. The randomisation is genuinely good, the setup takes minutes, and the statistical power is strong at scale. The problem is who is holding the pen. The platform is measuring its own effectiveness, using its own conversion data, with its own definition of the counterfactual, and it has an obvious commercial interest in the answer. Useful as a directional read, not as the number you rebuild a budget around. It also needs serious volume, generally in the region of hundreds of thousands of users per group, which rules out most brands below eight figures.
Geo holdout
Split the country into matched groups of markets. Keep spending normally in the control group. Turn the channel off, or scale it back sharply, in the test group. Compare revenue in your own Shopify data, not the platform's. Both groups experience the same season, the same promotion, the same competitor, the same weather. Whatever moves the market moves both sides equally, so what is left is the channel. It is slower and more fiddly than the other two, and it is the only one of the three where you own the measurement end to end.
How to design a geo holdout that survives scrutiny
Six decisions, made before anything is switched off. Get these right and the result will hold up when your CFO or your board pushes back on it.
Pick the one question
A geo test answers a single question. Is my Meta prospecting incremental? Is branded search buying me anything? Is my TikTok spend creating demand or just riding it? Do not attempt to test three channels at once in one design. You will end up unable to attribute the result to any of them, which is exactly the problem you started with.
Build matched market groups, not a random split
You want ten to fifteen geographic units at minimum, split into a test group and a control group that have tracked each other closely over the previous six to twelve months. Match on revenue level, growth trend, seasonality shape and customer mix, not on population. Two regions with similar populations and very different buying patterns are not a matched pair. If your data is thin, run the matching on weekly revenue over a year and look at the correlation, not the averages.
Set the minimum detectable effect from your margin
Before you run anything, work out how big a lift would actually change a decision. If a 5 percent difference in incremental revenue would not change your budget, you do not need to detect 5 percent, and pretending you can will send you chasing a test size you cannot afford. Most DTC brands are honest at a minimum detectable effect somewhere between 10 and 20 percent. Smaller markets and shorter windows push that number up.
Choose the window from your purchase cycle
Four to eight weeks for most DTC categories. You need one full consideration cycle plus a buffer, because demand you suppressed in week one can still convert in week four. Consumable and replenishment categories can often run at the shorter end. Considered purchases with a long research phase need the longer end, and a pre-period of two weeks to establish the baseline before anything changes.
Decide between full off and scale back
Going fully dark in the test markets gives the cleanest signal and the largest commercial risk. Scaling back by 50 to 70 percent is the usual compromise, and it still produces a readable result as long as the reduction is large enough to clear the noise. What you cannot do is quietly restore spend halfway through because the numbers got uncomfortable. That decision has to be made at the design stage and written down.
Freeze everything else
No new creative, no promotion changes, no landing page tests, no email campaign targeted at one group, no influencer drop in a test city. Every one of those is a confounder sitting inside your result. The test window is a boring window on purpose. Tell the team it is boring on purpose, because otherwise someone will helpfully launch something.
The five mistakes that void the result
Ending the test early is the most common one by a distance. Two weeks in, the test markets look flat, someone decides the ads clearly do nothing, and the decision gets made before the delayed conversions land. The opposite failure is just as frequent: a scary first week, spend gets restored, and the test is now a very expensive nothing.
Second, markets that are too small. A region doing 4,000 GBP a week has too much natural variance for a 15 percent effect to be visible. You will read noise as signal in whichever direction suits your prior.
Third, leaking channels. If you pause Meta but leave Google, TikTok, email and affiliate running at normal levels in both groups, that is correct and intended. If you pause Meta in the test markets and your agency simultaneously pushes Google harder there to protect the revenue, you have measured nothing except your agency's reflexes. Brief them explicitly.
Fourth, measuring the result inside the platform. The whole point is to escape the platform's view of the world. Your source of truth is Shopify revenue by region, pulled the same way for both groups, over the same window.
Fifth, treating one test as a permanent law. Incrementality decays and recovers. A channel that tested at 0.8x while you were running tired creative can test at 2.1x six months later on a new concept. Vendors who run these programmes properly re-test the major channels two to four times a year rather than once.
What to do with the number
A single iROAS figure is not a verdict on a channel. It is a calibration factor. If Meta prospecting reports 3.1x and tests at 1.6x, you now have a multiplier of roughly 0.52 to apply to that channel's reported performance until the next test. Do the same for every channel you test, and the platform dashboards become usable again, because you finally know how much to discount each one.
The reallocation follows from comparing calibrated returns rather than raw ones. The brand that discovers branded search is running below 1x does not necessarily switch it off, because there is usually a defensive argument about competitor conquesting. It does stop counting that revenue as acquisition, which changes the blended picture immediately and often reveals that a channel everyone had written off is carrying more of the real growth than anyone thought.
The mature version of this is a triangulation. A marketing mix model owns the strategic budget split across the portfolio, periodic geo tests calibrate the model so it does not drift, and platform attribution drops down to what it is actually good at, which is tactical optimisation inside a channel. Brands running that stack properly tend to report efficiency gains in the region of 10 to 30 percent in the first year, and most of that comes from stopping spend that was never doing anything rather than from finding a clever new channel.
What this looks like in practice
A wellness brand I work with was running a large retargeting budget that reported a 6.4x return. It was the best performing line in the account and nobody had questioned it in eighteen months. We built matched market groups on twelve months of Shopify revenue by region, scaled retargeting back by 70 percent in the test group, and left everything else alone for six weeks.
Week one looked alarming and the founder wanted to stop. We had agreed in advance that we would not, which is the only reason the test survived. By week six the gap between the groups had closed to a fraction of what the platform numbers implied. The incremental return on that retargeting line came out far below the reported figure, and a meaningful share of the budget was funding conversions that were already on their way.
We moved that budget into prospecting and into the post purchase flows, and the blended picture improved within the quarter. Reported ROAS on the account went down. Contribution profit went up. That is almost always the shape of a successful incrementality programme, and it is why these tests are politically difficult inside brands that reward the dashboard number.
Inside the system
How we build this for brands
The reason most brands never run a proper geo test is not that the statistics are hard. It is that nobody has the hours to build matched market groups from twelve months of order data, monitor a six week window without touching anything, and then reconcile the result against ad spend across four platforms. We run that work through a reporting agent built on live Shopify and ad data, which assembles the regional baselines, watches the test and control groups weekly, and flags the moment a confounder shows up in either group. The same profit and cash flow dashboards then carry the calibration factors forward, so every channel report the founder reads is already discounted to the tested reality rather than the platform's.
Around that sits the rest of the growth system: lifecycle and owned audience flows deployed in Klaviyo so the retention side is not silently carrying the acquisition numbers, and a VOC engine that mines reviews and support messages into the creative concepts we actually want to test next. Part of this runs live for portfolio brands today; the full system is what we deploy when we take a brand on.
Growth Audit
Find Out What Your Ad Spend Is Actually Buying
I will go through your channel reporting, show you where the platform numbers are almost certainly inflated, and design the first test worth running on your account. You get the test plan and the market groups, not a pitch deck.
Book Your AuditFrequently asked questions
What is incrementality testing in ecommerce?
Incrementality testing measures the sales that would not have happened without a given piece of marketing spend. You deliberately withhold spend from one group of customers or one set of geographic markets, keep running it everywhere else, and compare revenue between the two. The difference is the incremental contribution. Standard platform attribution cannot do this because it only counts conversions it can see and claim, not conversions it caused.
How much do ad platforms over-report conversions?
Independent measurement vendors consistently find that Meta and Google over-attribute by roughly 20 to 60 percent depending on channel and campaign type. Branded search is the worst offender because it largely cannibalises organic clicks the brand would have received anyway. One vendor reported a median incremental ROAS of 0.70x on branded search across 225 DTC geo tests.
How long should a geo holdout test run?
Four to eight weeks is the working range for most DTC brands. You need to cover at least one full purchase consideration cycle plus a buffer, because ad effects decay over days and weeks rather than stopping the moment you pause spend. Anything under two weeks is almost always too short to capture delayed conversions, and ending early is the single most common way operators void their own test.
How much ad spend do you need to run an incrementality test?
The constraint is statistical power rather than budget. Your test markets need enough baseline weekly revenue for a meaningful lift to be detectable above normal week to week noise. Most vendors point brands spending over 50,000 GBP a month per channel towards geo testing. Below that level, a platform conversion lift test or a single market scale back is usually the practical option.
What is the difference between incrementality testing and marketing mix modelling?
Marketing mix modelling uses aggregated weekly spend and sales data to estimate how each channel contributes across the whole portfolio. It is broad but correlational. Incrementality testing is narrow but causal, because you actually withhold spend and observe what happens. Use both: the model owns the strategic budget split, and periodic geo tests calibrate the model so it does not drift away from reality.
Why is a time based on off test not a valid incrementality test?
Turning ads off for two weeks and comparing revenue to the prior two weeks is the most commonly run and least reliable design. Seasonality, promotions, competitor activity, PR, weather and paydays all change between periods, and there is no control group to absorb any of it. A geo holdout solves this by running both conditions at the same time, so anything that moves the market moves both groups.
About the author
Caner Veli is a DTC operator who has helped 350+ brands fix broken growth engines. He built Liquiproof from zero to 3,000+ global retailers in under 6 years. He now runs the same playbook, supported by AI systems he built himself, for DTC and CPG brands.