Back to Store Roas Blog

Incrementality Testing for Ecommerce: From Holdouts to Geo Experiments

Incrementality testing in ecommerce answers a causal question: what extra revenue, profit, or customer value did a channel or tactic create that would not have happened anyway? That is different from attribution, which assigns credit to observed conversions. If you need to decide whether to keep, cut, or scale spend, incrementality is the right lens.

For advanced growth teams, the job is not to find a perfect model. It is to choose the right experiment design for the decision, verify the measurement is instrumented correctly, check that volume and statistical power are sufficient, and avoid contaminated results that waste budget. If volume is too low or the test is underpowered, do not run it yet.

If you are also working through attribution logic, start with Attribution measurement guide and AT-01. If your decisions are margin-led, pair this with Ecommerce profit analytics.

Start with the causal question, not the channel

A useful incrementality test starts with one clear sentence:

“If we reduce or pause spend on X for a defined population, period, or geography, what change occurs in incremental contribution profit, revenue, or qualified new customers?”

That framing matters because different teams need different outcomes:

Shopify’s KPI guidance supports this approach: choose metrics based on the business goal, not because they are easy to report (Shopify — Essential Ecommerce KPIs). Their ecommerce metrics guidance also stresses that different metrics belong on different cadences and should be owned by the right decision-maker (Shopify — Ecommerce Metrics).

Define the metric precisely

Do not test on a vague “ROAS lift”. Define the metric in operational terms before launch.

Use one primary metric and a small set of supporting metrics:

If your event instrumentation is wrong, the test is compromised before it starts. GA4 ecommerce reporting depends on correctly implemented ecommerce events and required parameters (Google Analytics — Ecommerce Purchases). In practice, that means confirming tracking completeness before you interpret any lift.

Pick the experiment design that fits the decision

There is no single best design. Choose the smallest design that can answer the causal question with acceptable risk.

Design Best for What is randomized Main limitation Typical failure mode
User holdout Email, SMS, onsite personalisation, some paid audiences Users or customer IDs Spillover and identity matching Contamination between exposed and control users
Audience holdout CRM or platform audiences Audience segments Small sample or segment drift Treatment leakage into control
Geo experiment Paid media, upper-funnel, markets with stable geography Regions, DMAs, countries Fewer units, more noise Weak power and seasonality imbalance
Time-based pause Simple tactical check Time periods Hard to separate test effect from trend/seasonality False lift or false decline from baseline drift

When holdouts are a good fit

Holdouts work well when you can reliably separate exposed and unexposed users, and when the unit of randomisation matches the decision. They are common for CRM, onsite experiments, and some ad platforms.

Use a holdout when:
- you can target distinct users or accounts,
- spillover is limited,
- and you need a clean user-level causal answer.

Do not use a holdout if the control group will be heavily contaminated by cross-device browsing, shared households, overlapping audiences, or manual campaign exclusions that the platform does not actually respect.

When geo tests are a better fit

Geo experiments are often better for paid media when user-level randomisation is not reliable or when the channel affects demand more broadly. They work by assigning regions to treatment and control, then comparing changes over time.

Use a geo test when:
- spend is large enough to support region-level analysis,
- markets are reasonably comparable,
- and the channel is likely to influence both direct and indirect demand.

Geo tests can be powerful, but only when you can manage seasonality, local differences, and interference across borders. They are less suitable for low-volume brands with too few markets.

Check sample size and power before you test

This is the key guardrail: if the volume is too low or the test is underpowered, do not run it.

Power is the probability of detecting a real effect of a chosen size. If your test cannot detect a lift large enough to matter commercially, the result may be inconclusive even if the tactic is working.

A practical sizing framework

For a simple two-group test, a rough sample size estimate for each group is:

n ≈ 2 × (Zα/2 + Zβ)² × σ² / Δ²

Where:
- n = sample per group
- Zα/2 = critical value for your confidence level
- = critical value for your desired power
- σ = standard deviation of the outcome
- Δ = minimum detectable effect, in the same units as the outcome

If you are using a conversion rate outcome, a simpler approximation is:

n ≈ 2 × (Zα/2 + Zβ)² × p(1-p) / Δ²

Where:
- p = baseline conversion rate
- Δ = absolute change in conversion rate you want to detect

Worked example

Assume:
- baseline conversion rate = 3.0%
- you want 80% power
- you use 95% confidence
- you want to detect a 0.6 percentage-point lift, from 3.0% to 3.6%

Using the proportion formula:
- Zα/2 = 1.96
- Zβ = 0.84
- p = 0.03
- Δ = 0.006

n ≈ 2 × (1.96 + 0.84)² × 0.03 × 0.97 / 0.006²

n ≈ 2 × 7.84 × 0.0291 / 0.000036

n ≈ 12,700 per group approximately

That is a planning estimate, not a guarantee. The true requirement may be higher if variance is larger, traffic is uneven, or the experiment has contamination.

Pre-test feasibility checklist

Before launch, confirm:
- the metric is instrumented correctly,
- the unit of randomisation matches the business question,
- the control group will remain isolated,
- the expected sample can reach the required size in a reasonable timeframe,
- and the minimum detectable effect is commercially meaningful.

If you cannot clear these checks, wait, redesign, or choose a different method.

Run the test safely

The mechanics matter. A poorly executed test creates false confidence.

Operational rules for the run phase

  1. Freeze the analysis plan
    - Define the primary metric, secondary metrics, duration, and stopping rule before launch.
    - Do not change the outcome after you see early results.

  2. Monitor for contamination
    - Check that users or regions assigned to control are not receiving treatment.
    - Watch for coupon codes, retargeting, shared lists, and internal traffic leakage.

  3. Track guardrail metrics
    - Revenue alone is not enough.
    - Monitor profit, refunds, repeat rate, and customer support signals where relevant.

  4. Keep a decision log
    - Record start and end dates, budget, audience rules, exclusions, and interruptions.
    - Shopify recommends combining quantitative metrics with qualitative feedback, so note customer complaints, ad fatigue, or fulfilment issues alongside the numbers (Shopify — Ecommerce Analytics Tools).

  5. Do not peek and react
    - Repeated early stopping without a formal plan inflates false positives.
    - If you need interim monitoring, define it as operational safety, not significance testing.

Common run-time risks

Read lift as a business decision, not a model output

The result should answer one question: should we change spend or strategy?

What to measure in the readout

At minimum, review:
- incremental revenue,
- incremental contribution profit,
- confidence interval around lift,
- cost of treatment,
- and a few guardrail metrics such as conversion rate, AOV, repeat purchase rate, or refund rate.

A positive lift is not automatically a good decision. If a campaign increases revenue but destroys margin, raises discount dependency, or attracts poor-quality customers, it may still be a bad investment.

That is why incrementality and attribution are different:
- Attribution says which touchpoints are associated with conversions.
- Incrementality says whether the activity caused extra outcomes.

Competitor positioning across the market reflects this split: some tools emphasize attribution and BI, while others combine MMM, creative analytics, and incrementality. The right choice depends on whether your decision need is tactical, portfolio-level, or financial (Triple Whale, ThoughtMetric, ShelfMerge, Nummbas).

Interpretation errors to avoid

A practical decision selector

Use this rule of thumb:

That is the difference between a useful incrementality test and an expensive reporting exercise.

Design your next test with the decision in mind

If you already trust your event tracking, the next step is to turn one commercial decision into a controlled experiment: pause, hold back, or split exposure in a way that can answer a profit-aware causal question.

If you need help framing the test, start with your primary metric, define the smallest effect worth acting on, and check whether the available volume can support a valid holdout or geo design. Then document the analysis plan before launch.

CTA: Design an incrementality test — define the causal question, choose the right design, and verify the volume, power, and contamination checks before you spend.