Online Experimentation - Index

Routing Summary

Statistics of large-scale online A/B testing: variance reduction, sequential / always-valid inference, trustworthiness checks, and interference-robust designs. Anchored by Deng et al. (2013, CUPED), Johari, Pekelis & Walsh (2015, always-valid inference), Howard et al. (2021, confidence sequences), Fabijan et al. (2019, SRM), Kohavi et al. (2012, puzzling outcomes), Johari et al. (2020, two-sided platforms), Bojinov et al. (2020, switchbacks) and the Larsen et al. (2022) review.

Concept Map

ConceptNoteTypeDepends OnKey Result
OCE framework and failure modesOnline Experimentation - OverviewoverviewThe Experimental Ideal; Potential Outcomes; Power Analysis makes 0.02% effects undetectable; four failure modes map to the notes below
CUPED / control variatesCUPED and Regression-Adjusted Variance ReductionmethodOverview; Power Analysis with ; about 50% reduction at Bing; never use post-treatment covariates
Peeking / optional stoppingThe Peeking Problem and Optional StoppingconceptOverview; Power Analysis; by the LIL a patient peeker rejects w.p. 1; 10 looks give about 20% FPR
Always-valid -values, mSPRTAlways-Valid p-values and the mSPRTmethodPeekingThm 1 duality with power-one sequential tests; ; Thm 2 first-order efficiency; Thm 3 mixing variance matches effect prior; commutes with Bonferroni and BH-G
Confidence sequencesConfidence SequencesconceptPeeking; mSPRT; Potential Outcomes; stitched LIL boundary; normal mixture ; empirical-Bernstein and sequential ATE; under CLT width
SRM and trust checksSample Ratio Mismatch and Trustworthiness ChecksapplicationOverview; The Experimental Ideal test on arm counts; about 6% of Microsoft experiments; five-stage taxonomy, ten diagnostic rules; A/A tests, carryover, OEC pitfalls; sequential multinomial test
Interference, CR / LR / TSR, clustersInterference and Marketplace ExperimentsconceptOverview; Potential OutcomesGTE estimand; CR unbiased when demand-constrained, LR when supply-constrained; naive estimators overestimate positive effects; TSR interpolates; cluster randomization trades bias for variance
Switchback experimentsSwitchback Experiment Design and AnalysismethodInterference; Randomization Inference-carryover; Horvitz–Thompson estimator unbiased; fair coins, flip every periods: ; exact FRT and conservative CLT

Notes

  • Online Experimentation - Overview — CONTAINS: OCE definition and notation (variants, units, metrics, OEC, guardrails), ATE and lift, table of four failure modes, Larsen sample-size example (147,456 vs 9.2 billion), triggered analysis, CV result for Sessions/User, early-trend fallacy (67% / 55%), relevance to geo experiments, ad auctions, lift studies and ABMs, pipeline pseudo-code.
  • CUPED and Regression-Adjusted Variance Reduction — CONTAINS: stratification variance decomposition, control variate estimator, optimal and theorem, unbiasedness of under randomization, stratification as special case, covariate choice, pre-period length (correlation vs coverage), missing pre-period data, pre-trigger covariates, delta method for page-level metrics, DQ-per-user post-treatment bias example, relation to ANCOVA and semiparametric efficiency, Bing slowdown experiment, Python sketch.
  • The Peeking Problem and Optional Stopping — CONTAINS: decision rule / sequential test / fixed-horizon -value definitions, union-over-looks argument, LIL “sampling to a foregone conclusion”, behavioural amplifiers (trends, novelty, primacy), six remedies (discipline, group sequential, SPRT with , thresholds, always-valid, Bayesian, bandits), costs of optional stopping, simulation table of false positive rates by monitoring schedule.
  • Always-Valid p-values and the mSPRT — CONTAINS: Definitions 1-2, Theorem 1 duality and proof sketch, mSPRT definition, Ville/martingale justification of the threshold, power one, Gaussian closed form and link to the normal mixture boundary, user model, aggressive / conservative / Goldilocks regimes, Theorems 2-3 and Eq. 12 for , robustness to mixing misspecification, Proposition 4 vs fixed horizon, Optimizely data, two-stream normal and Bernoulli tests, Bonferroni / BH-G / FCR results, limitations, simulation and code.
  • Confidence Sequences — CONTAINS: properties P1-P4, definition, Lemma 3 equivalences, sub- condition and uniform boundaries, linear boundary (Lemma 1), stitched boundary and finite LIL bound (Eq. 2), method of mixtures (Lemma 2), two-sided normal mixture (Eq. 14) and tuning of (Prop. 3), vs trade-off, empirical-Bernstein Theorem 4, sequential ATE with AIPW pseudo-outcomes (Corollary 2), running intersection caveat, width-ratio table, Python implementation.
  • Sample Ratio Mismatch and Trustworthiness Checks — CONTAINS: SRM definition and test, selection-bias argument with survival indicator , MSN carousel sign reversal, five-stage taxonomy table with examples and prevention, positive-cause SRMs, ten rules of thumb, A/A tests, carryover (3 weeks to 3+ months), OEC decomposition, Lindon–Malek sequential multinomial Bayes-factor test, fixed and sequential code examples, lift-study and geo readings.
  • Interference and Marketplace Experiments — CONTAINS: SUTVA / interference / GTE definitions, network vs marketplace interference, reported bias magnitudes, Markov chain model and mean-field ODE (Eqs. 7-8), market balance , CR / LR / TSR designs and naive estimators (Eqs. 19-21), Theorems 3-4 and Proposition 4 with intuition, TSR allocation rule (Eq. 26), TSRI estimators, bias-variance simulations, graph-cluster and ego-cluster randomization, budget-split designs, geo experiments as cluster designs, lodging and advertising examples, simulation as a design tool.
  • Switchback Experiment Design and Analysis — CONTAINS: non-anticipation and -carryover assumptions, lag- estimand, regular switchback definition, Horvitz–Thompson estimator, minimax design problem, Theorems 1-2 and the , example, granularity and robustness corollaries, exact randomization test (Algorithm 1), variance bound , Theorem 3 CLT, misspecified , procedure for identifying , horizon planning, Python design / estimator / test code, simulation comparing HT with the naive contrast, media-flighting interpretation.

Sources

  • Deng 2013 - CUPED Improving Sensitivity with Pre-Experiment Data — Deng, A., Xu, Y., Kohavi, R. & Walker, T. (2013), “Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data,” WSDM 2013.
  • Johari 2015 - Always Valid Inference — Johari, R., Pekelis, L. & Walsh, D. (2015; v3 2019), “Always Valid Inference: Continuous Monitoring of A/B Tests,” arXiv:1512.04922 (Operations Research 2022).
  • Howard 2021 - Time-uniform Nonparametric Confidence Sequences — Howard, S., Ramdas, A., McAuliffe, J. & Sekhon, J. (2021), “Time-uniform, nonparametric, nonasymptotic confidence sequences,” Annals of Statistics; arXiv:1810.08240.
  • Larsen 2022 - Statistical Challenges in Online Controlled Experiments — Larsen, N., Stallrich, J., Sengupta, S., Deng, A., Kohavi, R. & Stevens, N. (2022), “Statistical Challenges in Online Controlled Experiments: A Review of A/B Testing Methodology,” arXiv:2212.11366.
  • Kohavi 2012 - Trustworthy Online Controlled Experiments Five Puzzling Outcomes — Kohavi, R., Deng, A., Frasca, B., Longbotham, R., Walker, T. & Xu, Y. (2012), “Trustworthy Online Controlled Experiments: Five Puzzling Outcomes Explained,” KDD 2012.
  • Fabijan 2019 - Diagnosing Sample Ratio Mismatch — Fabijan, A. et al. (2019), “Diagnosing Sample Ratio Mismatch in Online Controlled Experiments: A Taxonomy and Rules of Thumb for Practitioners,” KDD 2019.
  • Lindon 2020 - Anytime-Valid Inference for Multinomial Count Data — Lindon, M. & Malek, A. (2020), “Anytime-Valid Inference for Multinomial Count Data,” arXiv:2011.03567.
  • Johari 2020 - Experimental Design in Two-Sided Platforms — Johari, R., Li, H., Liskovich, I. & Weintraub, G. (2020), “Experimental Design in Two-Sided Platforms: An Analysis of Bias,” arXiv:2002.05670 (Management Science 2022).
  • Bojinov 2020 - Design and Analysis of Switchback Experiments — Bojinov, I., Simchi-Levi, D. & Zhao, J. (2020), “Design and Analysis of Switchback Experiments,” arXiv:2009.00148 (Management Science 2023).