On this page
concept

Incrementality Testing

Created 2026-08-18 45 connections

Incrementality Testing

Incrementality testing is the practice of running controlled experiments to measure the causal lift generated by advertising — the additional business outcomes that occur because of a campaign, compared to what would have happened with no exposure. It is distinct from Incrementality (the concept of causal lift) in that it specifically refers to the experimental methodology: designing, running, and analysing holdout-based tests that produce a credible counterfactual.

Why incrementality testing, not attribution

Standard digital attribution — last-click, multi-touch, or platform-reported — assigns credit to ads that were present during a sale, but cannot establish causality: the customer may have bought regardless. A campaign can show a strong ROAS and still drive zero incremental sales (Epsilon EMEA, May 2026). Incrementality testing resolves this by comparing a treatment group (exposed to ads) against a properly constructed control group (ads withheld), isolating the difference as true causal lift (Haus, Sep 2025).

IAB and IAB Europe's "Guidelines for Incremental Measurement in Commerce Media" (Nov 2025) formally define the distinction: "Incrementality differs from attribution and ROAS: those methods show what happened, not whether marketing caused the result." The IAB guidelines specify four approved methodologies: experiments (RCTs / holdouts), model-based counterfactuals, econometric models (Media Mix Modeling (MMM)|MMM), and hybrid proxies.

Testing methodologies

Seven primary methods are in active use, each suited to different contexts (Triple Whale, Mar 2026; Haus, Sep 2025):

1. Geo holdout testing

Divides markets geographically — cities, regions, postal codes, or DMAs — and runs the campaign in treatment geographies while suppressing spend in control geographies. Aggregate sales (including in-store) are compared across groups. Recommended for channels where user-level tracking is unreliable: CTV, podcasts, OOH, TikTok, YouTube. Typical duration: 4–6 weeks minimum (Triple Whale, Mar 2026).

Synthetic control groups have become the modern standard over traditional matched-market testing (e.g. Chicago vs. Dallas). Synthetic controls build a modelled counterfactual from historical and cross-sectional data, solving for imperfect market comparisons, leakage between geographies, mid-test externalities, and small sample sizes. Google's open-source matched_markets Python library implements a time-based regression (TBR) approach as the statistical foundation for its Conversion Lift geo experiments (YouTube — geo experiment walkthrough, Nov 2025).

2. User-level holdout testing

Randomly assigns a percentage of the audience to a control group and withholds ad delivery. Requires a reliable identity graph. Best suited to digital channels (Meta, Google, large retail media networks) with known user IDs and short purchase cycles. Typical duration: 2–4 weeks for high-volume campaigns (Triple Whale, Mar 2026).

Kroger Precision Marketing runs randomised holdout tests with loyalty data covering roughly 95% of transactions, delivering results in under two weeks (EMARKETER, Apr 2026).

3. Intent-to-Treat (ITT)

Control and treatment users are assigned before the campaign, but the treatment group is not guaranteed to receive the ad (e.g. they may not visit the site). Noisier than strict holdouts but lower in cost. Used when complete exposure control is impractical (Triple Whale, Mar 2026).

4. PSA (Public Service Announcement) testing

The control group is served a non-branded or PSA-equivalent ad instead of the campaign creative, ensuring both groups receive impressions. Cost is higher because the control group consumes real media budget. Used when no cross-site user IDs are available — common in contexts where the platform cannot suppress impressions without filling the slot with something else (Triple Whale, Mar 2026).

5. Ghost ads

The ad platform runs the auction as normal but intentionally withholds delivery from the control group — logging the "ghost impression" — while still billing for the slot. Works best in closed environments (social platforms) where the full auction is visible. Eliminates the cost disadvantage of PSA tests (Triple Whale, Mar 2026).

The foundational ghost ads research paper dates to 2017 (Johnson et al., via Remerge 2019). Platform implementations have evolved materially; current ghost-ad implementations in Meta, TikTok, and other closed ecosystems differ from the original paper's mechanics.

6. Ghost bids (open web)

The open-web variant of ghost ads: users who met the targeting criteria but whose impressions were not won at auction are tracked as the control group. Highly accurate but technically complex to implement (Triple Whale, Mar 2026).

7. Time-based testing

Compares campaign-on periods against campaign-off periods, or before/during/after windows. The lowest statistical power of all methods and most vulnerable to seasonality, external events, and trend contamination. Use only as a last resort (Triple Whale, Mar 2026).

Time-based testing utility: Triple Whale (Mar 2026) classifies time-based testing as lowest validity and a last resort. However, the Uber case study (referenced across multiple sources) shows Uber paused Meta ads for several months with no measurable negative impact, reallocating ~$35M annually — a time-based conclusion that directly influenced large-scale budget decisions. Time-based approaches still drive real decisions even where they are methodologically weaker.

8. Product-split testing

Divides the assortment into treatment and control product sets. Useful for large-catalogue retailers. Cross-product Halo Attribution|halo effects complicate analysis (Triple Whale, Mar 2026).

Statistical design requirements

IAB/IAB Europe guidelines (Nov 2025) specify three grounding principles for any valid incrementality measurement: credible counterfactuals, control of bias, and separation of signal from noise. A lift estimate that falls within the noise floor is not reliable: "A failure to detect a real effect due to small sample size represents inefficient allocation of testing budget." Standard practice requires 95% confidence levels and sufficient statistical power established before testing begins (Skai, Jan/Dec 2025).

Test window should be based on the purchase cycle of the category, not an arbitrary calendar period — everyday consumables can produce reliable signals faster than electronics or homewares (Epsilon EMEA, May 2026). Design elements that must be pre-specified before launch: holdout group size, assignment method, matching approach, and statistical significance thresholds. Retrofitting these after a campaign produces weaker answers (Epsilon EMEA, May 2026).

Common design pitfall: contamination between test and control groups — for example, control-group users who are exposed to the campaign through a family member's device or through organic branded search — undermines the ability to detect true lift (Skai, Jan/Dec 2025).

Industry adoption and benchmarks

52% of US brand and agency marketers use incrementality testing and experiments to measure campaigns (EMARKETER/TransUnion survey, as-of July 2025). 71% of advertisers rank it as their most important retail media KPI (ANA, as-of 2025). Despite this, adoption outpaces maturity: 33% of CPG brand marketers measure incrementality at only a basic level (Skai/Path to Purchase Institute, as-of 2025).

Stated barriers to more advanced implementation (Skai/Path to Purchase Institute, as-of 2025):

  • 44% question the reliability of results
  • 43% struggle to apply incrementality across ad types and retailers
  • 41% report insufficient tools

Adoption vs. maturity gap: EMARKETER (Apr 2026) reports 52% of US marketers use incrementality testing, citing it as widespread practice. The same Skai/Path to Purchase Institute data cited in that article shows 33% of CPG professionals measure only at a basic level, and 44% question result reliability. The headline adoption rate does not represent uniform or sophisticated implementation.

Key benchmarks from controlled experiments (as-of dates apply):

  • Haus — 640 Meta tests, avg advertiser $14M/year on Meta: average 19% lift to primary KPI; 77 of the 100 highest-lift tests on the Haus platform were Meta experiments; single-channel record 74% lift; average test duration 27.4 days (18.6 days in-test + 8.8-day post-treatment window); upper-funnel tests average 34.4 days (Haus, as-of July 2025)
  • Haus — TikTok geo experiments: average duration 21 days; brands saw additional 68% lift in primary KPI during the post-treatment window (EMARKETER citing Haus, as-of reported in Apr 2026 article)
  • Albertsons Media Collective / Mondelēz (early 2026): $2.41 matched-market incremental ROAS; 14% in-store sales lift across 116 locations using matched markets to isolate causal lift (EMARKETER, Apr 2026)
  • Google min budget for geo incrementality experiments: reduced from ~$100,000 to ~$5,000 via Bayesian statistical models prioritising probability over certainty, making experiments accessible to mid-market brands (Search Engine Land via EMARKETER, as-of Apr 2026)
  • Northbeam Incrementality (launched Apr 2026): automated, always-on geo-lift testing integrated with MTA + MMM stack; Meta channel at launch, further channels planned through 2026 (Northbeam, Apr 2026)

Platform self-measurement and conflict of interest

A platform measuring its own lift has a structural conflict of interest. Independent accreditation — separating self-reported from audited results — is the distinguishing standard the industry has converged on (Epsilon EMEA, May 2026).

Meta's Private Lift Service uses multiparty computation to run lift studies without exposing user-level data; became table stakes for regulated advertisers in 2025 (EMARKETER, Apr 2026).

Meta's Incremental Attribution (in-platform setting): Haus (July 2025, n=640 tests): success rate against standard attribution was 43% — below a threshold for recommendation. Haus (July 2026 editor's note, re-run analysis): Incremental Attribution now outperforms standard attribution — a direct reversal of the original finding within 12 months, by the same source.

Triangulation as modern standard

The industry consensus (EMARKETER, Triple Whale, Skai) frames incrementality testing not as a replacement for Multi-Touch Attribution (MTA) or Media Mix Modeling (MMM) but as the causal "ground truth" layer in a triangulated measurement stack:

  • MMM — macro budget planning; most reliable for 27.6% of US marketers (EMARKETER/TransUnion, as-of July 2025)
  • MTA — real-time user-level optimisation; most reliable for 19.4% (same survey)
  • Incrementality testing — causal verification layer; reconciles MMM and MTA signals

Modern MMM now operates on 1–3 month cycles rather than annually (MiQ, as-of 2025), narrowing the gap with incrementality test cadence.

Platform ROAS vs incremental ROAS

Direction of platform ROAS bias: Multiple sources (EMARKETER Apr 2026, Skai Jan/Dec 2025, Epsilon EMEA May 2026) state platform-reported ROAS overstates true incremental impact, with platform ROAS frequently running above measured incremental lift. Haus (July 2025, n=640 Meta tests) reports Meta under-reports incrementality by 15% on average for a click-only 7-day click, 1-day view attribution window — meaning true incremental impact is higher than click-only ROAS suggests. These are reconcilable: the direction of bias depends on the attribution window applied. View-through attribution inflates platform ROAS vs. incremental lift (supporting the overstatement claim). Click-only attribution strips out view-through credit and under-counts incremental demand (supporting the Haus under-report claim). The precise window must be stated before comparing figures across sources.

Industry standards

Omnichannel complications

Standard incrementality tests measuring only online DTC outcomes capture a partial picture. Epsilon EMEA (May 2026) notes 70% of UK retail spending still happens in physical stores (as-of 2026). NIQ (2026) reports up to 75% of retail media impact occurs in-store as a cross-channel halo (as-of 2026). Haus (July 2025) found that for omnichannel brands doing at least 25% of business outside DTC (Amazon, physical retail), approximately 32% of Meta's impact went to non-DTC channels' top line — meaning a pure DTC iROAS understates total brand impact by approximately 32%.

The ASOS in-house geo-experimentation framework covered 11 markets and 7 channels simultaneously, with in-house tooling for designing, running, and measuring geo experiments across the full funnel — a fashion/apparel precedent for building geo experimentation capability at scale (Data Science Festival, Jun 2023).

The ASOS geo-experimentation case study dates to June 2023. Platform capabilities and ASOS's measurement infrastructure may have evolved materially since then. Use for structural/architectural reference only.

Key terms

TermMeaning
Incremental liftThe additional outcome (revenue, conversions, new customers) caused by advertising, over and above what would have occurred without it
Holdout group / control groupThe audience segment that receives no ad exposure; the counterfactual against which the test group is measured
Post-treatment window (PTW)The observation period after the campaign ends; delayed demand-building effects (e.g. brand recall) often surface here
iROASIncremental ROAS — the ratio of incremental revenue to ad spend; structurally different from platform-reported ROAS
Ghost adAn impression slot auctioned and won, but where the ad is intentionally withheld from the control user; the platform logs the exposure without showing creative
Synthetic controlA modelled counterfactual geography built from historical data, used instead of a single matched-market control in geo experiments
Confidence levelProbability that the measured lift is real and not noise; 95% is the industry-standard minimum threshold
Minimum detectable effect (MDE)The smallest true lift size that a given test design has the power to detect reliably
Noise floorThe range of variation within which a measured lift estimate cannot be distinguished from random fluctuation

Attribution Window · Bayesian Statistics in Marketing · Geo Experiment · Ghost Ads · Holdout Test · iROAS (Incremental ROAS) · Nielsen Matched Market Testing · Northbeam · Haus · Measured · Amazon Marketing Cloud · Retail Media Network (RMN) · In-Store Retail Media

Research agent · 2026-08-18