How does advertising carryover (adstock) violate the assumptions of switchback experiments, always-valid sequential tests and geo tests, and what design changes fix it?

Summary

Adstock is interference across time: today’s outcome depends on the whole past assignment path. Each design copes with it through a different assumption, and adstock strains each one differently. Switchbacks assume carryover ends after periods, but geometric adstock never ends, so must be a truncation horizon and the effective sample size collapses to coin flips. Always-valid tests keep their type I guarantee (no effect means nothing to carry over) but the estimand drifts: the effect wears in, so the earlier the test stops the more it understates the sustained effect, and geo time series are not i.i.d. streams. Geo tests need an adstock-free pretest, a cooldown that ends when the cumulative effect flattens, and separate estimands for heavy-up and go-dark. The fix is always the same trio: size every window from the retention rate, target the sustained-treatment estimand explicitly, and separate the stopping decision from the magnitude read-out.

Answer

Q - Carryover Dynamics and the Timing of Sequential Media Experiments established that read-out windows must span the adstock tail, that back-to-back tests contaminate each other, and that a state-space model can absorb carryover. This note extends it to the online-experimentation cluster: which formal assumption of each design does adstock break, how badly, and what should change?

1. One idea, three disguises: carryover is interference in time

Interference and Marketplace Experiments defines interference as any failure of and the decision-relevant estimand as the global treatment effect, everyone treated versus no-one treated. Switchback Experiment Design and Analysis removes cross-unit interference by treating the market as one unit and immediately meets the temporal version: potential outcomes are and the estimand is the lag- effect of sustained treatment, (^def-lag-p-effect). Time-Varying Treatments and G-computation writes the same object as for two treatment sequences. These are the same idea. Adstock supplies the mechanism: the regressor is with (geometric) or (delayed peak) (Carryover (Adstock) Functional Forms), the Koyck lag with long-run multiplier (^def-koyck).

Synthesis: under geometric adstock and a locally linear response, the effect periods after a sustained switch-on is : a wear-in curve. The running average over the first periods is , which for weekly is 56% of at , 73% at and 82% at . Everything below follows from this curve and its mirror image after switch-off.

What only looks similar: conversion delay (Delayed and Censored Feedback - Overview). Adstock delays the effect; conversion lag delays the measurement of an outcome that has already been caused. Both lengthen the effective , but they need different repairs: a longer window for the first, a censoring likelihood for the second.

2. What each design assumes and how adstock strains it

DesignAssumption that carries the inferenceHow adstock violates itSymptomRepair
Switchback (Bojinov et al.)-carryover: only the last assignments matter (Switchback Experiment Design and Analysis)Geometric weights are never zero; delayed adstock peaks after the switch; ratchet or historical-maximum response has unbounded memory (Carryover Effects and Distributed Lags) makes biased toward zero; honest leaves few coin flipsSet ; periods shorter than carryover; Fisher test for the sharp null; abandon the design if is small
mSPRT / always-valid -valuesi.i.d. draws from with fixed (The Peeking Problem and Optional Stopping) ramps up with wear-in; time-aggregated sales have MA(1) Koyck errorsNull control survives, but the estimate at the stopping time understates and omits the post-stop tailUse the test for sign or harm; pre-commit a minimum run and cooldown for magnitude
Confidence sequencesSub- supermartingale; estimand may vary, (Confidence Sequences)Drift makes the running intersection empty; the covered quantity is the running mean effect, not the sustained effectValid interval for the wrong estimandNon-intersected sequence with estimand ; predictable that includes lagged adstock
GBR geo testPretest is unaffected by treatment; test window captures the responseResidual adstock from earlier flights sits in the pretest; offline response lags by Biased covariate; truncated ROASWashout before pretest; extend test period by
TBR geo testStable regression of the treatment aggregate on the control aggregate, absent the intervention (TBR Design Sensitivity and the Stationarity Assumption)Decaying pre-test adstock in treatment geos is a differential trend; effect continues into cooldownCounterfactual drifts; has not flattenedCooldown until flattens, no longer; BSTS or TBR-OR under sustained trends

3. Switchbacks: the truncation horizon eats the sample size

The Horvitz–Thompson estimator uses only periods whose whole window is all-treated or all-control, a data-driven washout (^thm-ht-unbiased). If the exact test remains valid for the sharp null but is biased for ; if everything stays valid, only less efficient. The optimal design flips a fair coin once every periods (^thm-optimal-switchback), so the sample size is coin flips, not users or weeks.

Synthesis: for geometric adstock there is no true . With linear response, , the lower end applying when assignments before the window are independent of it. So pick for a tolerated relative bias . The vault’s benchmarks: Jin et al. treat weeks as effectively infinite for , and the 90% duration interval for packaged goods is 6 to 9 months (^thm-duration). The switchback CLT needs and the test that identifies needs , so a 13-week horizon asks for decades of data.

Own simulation (not from the papers): weekly switchback, , four years

, , , so and weeks. Outcomes centred on the known baseline; 4,000 assignment paths per design.

Design Coin flipsHT mean HT sdNaive diff. in means Naive sd
0 (flip weekly)2080.300.280.300.16
21020.710.730.400.18
4500.891.050.560.21
8240.971.470.730.22

The HT means track the bound (0.30, 0.66, 0.83, 0.96). Nearly unbiased estimation () costs a standard deviation of 44% of the effect after four years; weekly on/off pulsing recovers only 30% of it. The naive contrast is tight and wrong at every block length.

Two further violations are qualitative. Ratchet and historical-maximum models make sales depend on , so no finite exists and on→off is not the mirror image of off→on. And if spend follows a rule that reacts to lagged sales, as Design of Dynamic Response Models says is typical, past outcomes become time-varying confounders affected by past treatment; either the experimenter, not the pacing algorithm, must own the coin, or the analysis needs the g-formula under sequential ignorability.

Switchbacks therefore suit responses with almost no memory. In Vaver and Koehler’s search experiment incremental clicks flattened the instant spend returned to baseline () while offline sales kept climbing (Geo-Experiment Design and Power Analysis): the first metric is switchback material, the second is not.

4. Always-valid tests: the error rate survives, the estimand does not

Synthesis: carryover of the treatment cannot inflate type I error under the strict null, because a zero effect has nothing to carry over; Ville’s inequality is applied under (^def-msprt). Three other things go wrong.

  1. Stopping early means stopping on the steep part of the wear-in curve. The mSPRT’s selling point is that large effects stop very early. Under adstock the early effect is only , and ^roas-eq requires the numerator to run to . An estimate frozen at the stopping time is biased down by wear-in and tail truncation and up by stopping on a favourable fluctuation (the type M bias in the peeking note); the two do not reliably cancel.
  2. Null contamination from earlier flights. Residual adstock from a previous campaign, or users recycled from an earlier experiment, makes the arms differ at with no current treatment. Kohavi et al. report such carryover from re-used hash buckets lasting three weeks to more than three months (Online Experimentation - Overview). This does produce false positives. Run an A/A period, or wash out for .
  3. The data are not the stream the theory assumes. For user-level tests with fresh randomized arrivals, Howard et al.’s design-based sequential ATE needs neither independence nor a common mean and covers the running (^thm-empirical-bernstein). In a geo test randomization happened once; the daily treatment-minus-counterfactual series is autocorrelated (the Koyck error is MA(1)), and dependence is already listed among the mSPRT’s limitations. The vault has no always-valid method for that case (see Gaps).

Delayed conversions add a measurement layer. In Chapelle’s data only 35% of conversions arrive within an hour and 13% arrive after two weeks; labelling pending users as negatives biases the rate down, worst for the freshest cohort (Delayed Feedback Model for Conversion Prediction). Both arms are censored alike, so the null is safe. Synthesis: if the ad changes the delay distribution (purchase acceleration), early looks show a lift in “converted so far” that shrinks as control catches up. Fix it by modelling “ever” and “when” jointly (^thm-dfm-likelihood) or by the delay-corrected estimator (^def-bandit-delay-corrected-estimator). Vernade et al.’s result that delay alone costs nothing asymptotically while a hard window rescales every rate by assumes a delay CDF shared across arms, which acceleration breaks.

5. Geo tests: clean pretest, flat cooldown, one-sided estimands

Geo tests are cluster randomization against cross-unit interference (Interference and Marketplace Experiments); nothing in them protects against the temporal kind.

  • Pretest. GBR regresses test-period response on pretest response; TBR fits its counterfactual on pretest data. Both are adjustments on a pre-period covariate, so the one hard rule of CUPED and Regression-Adjusted Variance Reduction applies: never use a covariate the treatment could have affected. Adstock from an earlier flight in the same geos is such an effect. Synthesis: leave at least between the last spend change and the first pretest day, or the decaying stock becomes the differential trend that biases TBR and drops its coverage (TBR Design Sensitivity and the Stationarity Assumption).
  • Cooldown. TBR reads until it flattens; a curve that never flattens signals a long-lived lagged effect or a broken model. But at fixed total spend the iROAS half-width grows with cooldown length, roughly . Synthesis: each extra cooldown day adds signal of order and a full unit of noise variance, so stop at and report the truncated share as a stated bias rather than waiting for a perfectly flat line.
  • Go-dark versus heavy-up. A holdout removes spend from a stock built before the test; the effect wears out along , so a short holdout understates the channel just as a short heavy-up does. Under ratchet response () the two designs estimate different parameters and should not be pooled.
  • Estimate at the finest grain. Clarke’s aggregation bias makes carryover estimated from annual data 20 to 50 times longer than from monthly data (^thm-clarke).

Practical Implications

A decision rule for choosing and configuring the design:

  1. Compute the horizon. From an MMM posterior for at weekly or daily grain, with to . Use the upper posterior quantile of ; the cost of is variance, the cost of is bias.
  2. Count coin flips. . If is below a few dozen (my heuristic; the source only says the CLT needs large and the -test needs ), do not run a switchback; use geo replication. If it is large (search, retail-media, pricing), run the optimal design, analyse with Horvitz–Thompson and the Fisher randomization test, and never with a difference in means.
  3. Split the two questions. Monitor continuously with a confidence sequence for harm or sign; read magnitude once, at a pre-registered time treatment length plus , over the cumulative window of the ROAS definition.
  4. Wash out before, cool down after. Both sized by ; A/A-check the pretest.
  5. Name the estimand (sustained-on versus sustained-off, a finite flight, heavy-up or go-dark) and simulate the planned assignment path through the MMM or ABM first, as the interference note recommends for marketplaces, to get each estimator’s bias before spending.
  6. Handle conversion lag as censoring, not by waiting: a delay model for user-level tests, or a fixed attribution window applied identically to both arms and declared part of the estimand.

Source Notes

NoteRelevance
Switchback Experiment Design and Analysis-carryover, HT estimator, optimal switching, misspecified
Interference and Marketplace ExperimentsSUTVA, global treatment effect, simulation as a design tool
Always-Valid p-values and the mSPRT · The Peeking Problem and Optional Stoppingi.i.d. fixed- setup, stated limitations, type M bias at stopping
Confidence SequencesTime-varying estimand, empirical-Bernstein sequence, running intersection under drift
Online Experimentation - OverviewCarryover from re-used buckets lasting weeks to months
Carryover (Adstock) Functional Forms · Carryover Effects and Distributed Lags · Design of Dynamic Response ModelsAdstock forms, , Koyck MA(1) error, ratchet, aggregation bias, spend rules
Advertising and Promotion Effects90% duration interval of 6 to 9 months
Geo-Experiment Design and Power AnalysisGBR model, lag , clicks versus offline sales
Time-Based Regression Estimator for Geo Experiments · TBR Design Sensitivity and the Stationarity AssumptionCooldown, flattening of , stability, half-width versus cooldown
ROAS, mROAS, and Optimal Media MixROAS numerator runs to
Delayed and Censored Feedback - Overview · Delayed Feedback Model for Conversion Prediction · Bandit Models with Delayed and Censored FeedbackMeasurement delay, censoring likelihood, delay-corrected estimator
Time-Varying Treatments and G-computation · CUPED and Regression-Adjusted Variance ReductionSequence estimands; clean pre-period covariates
Bojinov 2020 - Design and Analysis of Switchback ExperimentsSecs. 2-4 and 6

Gaps

  • No always-valid inference for autocorrelated or regression-adjusted time series. The vault’s sequential notes cover i.i.d. streams and design-based sequential randomization; nothing covers anytime-valid monitoring of a TBR or BSTS counterfactual.
  • No switchback theory for infinite or model-based carryover. The bias bound and the simulation above are my own; the ingested paper treats only fixed finite . Regression-adjusted or model-assisted switchback estimators are mentioned in one sentence only.
  • No formal optimal cooldown or washout length (still open from the earlier Q&A), and no note on long-term holdouts or novelty-effect estimation.
  • Hysteresis and asymmetric response appear only as functional forms, with no experimental-design treatment.

Follow-Up Questions

  • Can a geometric-adstock outcome model be combined with the Horvitz–Thompson switchback estimator to de-bias short blocks while keeping randomization-based inference?
  • What is the variance-optimal cooldown length in TBR as a function of , and spend intensity?
  • How should a confidence sequence be built for the cumulative TBR effect with estimated regression parameters?