Confidence Sequences
Summary
A confidence sequence is a sequence of intervals with : the coverage statement holds simultaneously for all sample sizes, hence at any data-dependent stopping time. Howard, Ramdas, McAuliffe & Sekhon (Annals of Statistics 2021) build them from uniform boundaries for a centred sum measured against an “intrinsic time” , under a nonparametric sub- condition (a supermartingale bound on ). Two constructions matter in practice: the closed-form stitched boundary, growing at the law-of-the-iterated-logarithm rate , and conjugate mixture boundaries such as the normal mixture , which is exactly the Gaussian mSPRT made nonparametric. The cost of anytime validity is less than a doubling of the fixed-sample CLT width over five orders of magnitude of .
Overview
The paper is motivated directly by A/B testing: experiments “are inherently sequential”, results “are often monitored continuously using inferential methods that assume a fixed sample”, and most tests “are run with little formal planning and fluid decision-making” compared with clinical trials (Sec. 1). See The Peeking Problem and Optional Stopping. The authors want intervals with four properties:
- (P1) Nonasymptotic and nonparametric — coverage at every sample size without distributional assumptions.
- (P2) Unbounded sample size — no horizon needs to be fixed; one may tune for a planned but always keep sampling.
- (P3) Arbitrary stopping rules — no assumption on how the experimenter decides to stop or act.
- (P4) Asymptotically zero width — widths shrink at up to log factors.
Fixed-sample CLT intervals satisfy none of P1-P3. The idea dates to Darling & Robbins (1967); the same objects are called repeated confidence intervals (Jennison & Turnbull), always-valid confidence intervals (Johari et al.) and anytime confidence intervals in the bandit literature, where they drive best-arm identification.
Main Content
Confidence sequence (Eq. 1) ^def-confidence-sequence
For , a -confidence sequence for a (possibly time-varying) estimand is a sequence of sets , typically intervals , such that
Equivalent forms of time-uniformity (Lemma 3) ^thm-equivalence
For an adapted sequence of events the following are equivalent: (a) ; (b) for all random times , not necessarily stopping times; (c) for all stopping times . Taking shows the definition above coincides with Johari et al.’s stopping-time definition; taking shows an always-valid -value is equivalently one with .
So confidence sequences, always-valid -values and sequential tests are three views of one object: reject the first time . Whenever the radius this is a test of power one.
From tail bounds to sequences
Let be the running average conditional mean and . If for some adapted intrinsic time (a variance process; in the simplest case), then is a lower confidence sequence for . Apply the same to and take a union bound for two-sided sequences.
Sub- process and uniform boundary (Definitions 1-2)
is sub- with variance process if for each there is a supermartingale with and
A function is a sub- uniform boundary with crossing probability if over all sub- pairs .
plays the role of a cumulant generating function. The catalogue is sub-Gaussian , sub-Bernoulli, sub-Poisson, sub-exponential, and sub-gamma . Any process with an MGF near zero is sub-gamma (Prop. 1), so a sub-gamma boundary is a universal fallback.
Linear boundary (Lemma 1) ^thm-linear-boundary
For any ,
is a sub- uniform boundary with crossing probability .
This is Ville’s inequality applied to one exponential supermartingale; in the Gaussian case it is Wald’s SPRT line. Because , the resulting interval never shrinks to zero (P4 fails): a single is a single point alternative. Two ways to bend the line give the paper’s main tools.
Stitching (Theorem 1)
Split intrinsic time into geometric epochs , use a linear boundary tuned to each epoch with error budget where , and union-bound (“peeling”). With this yields the polynomial stitched boundary, a finite LIL bound, with arbitrarily close to 2. With , , , for i.i.d. 1-sub-Gaussian observations:
Corollary 1 recovers the classical upper LIL, , confirming the rate is unimprovable.
Conjugate mixtures (Lemma 2)
Integrate the exponential supermartingale against a distribution on : is again bounded by a supermartingale, and is a uniform boundary. This is Robbins’ method of mixtures, the same device as the mSPRT.
Two-sided normal mixture boundary (Eq. 14) ^def-normal-mixture
For a sub-Gaussian process, mixing over gives the closed form
so for i.i.d. observations with variance (proxy) the confidence sequence is . The tuning parameter sets where the boundary is tightest: is minimised at with (Prop. 3, the lower Lambert branch).
Every mixture with a density positive at the origin grows like (Prop. 2), slightly faster than the LIL rate. The authors argue this is the better practical trade: linear boundaries degrade fastest away from their optimised time, mixtures much more slowly, stitched boundaries slowest, and mixtures are tighter over the range one actually cares about (Sec. 3.5, Fig. 4). In the sub-Gaussian case mixture boundaries are unimprovable (Sec. 3.6). Analogous beta-binomial, gamma-Poisson and gamma-exponential mixtures cover the other families. In a parametric exponential family, rejecting when the mixture sequence excludes is the mSPRT (Sec. 6); confidence sequences are its nonparametric generalisation.
Empirical-Bernstein sequence and the sequential ATE
Empirical-Bernstein confidence sequence (Theorem 4) ^thm-empirical-bernstein
Suppose a.s., let be any -valued predictable sequence, and let be a sub-exponential uniform boundary with scale and crossing probability . Then
The intrinsic time is the observed sum of squared prediction errors, so the width adapts to the true variance (like a -test) with no variance knowledge, no common mean and no independence. Better predictions (trends, seasonality, regression on covariates, ML) shrink the interval, but coverage holds for any predictions. This is the sequential counterpart of CUPED.
Sequential ATE under the Neyman–Rubin model (Sec. 4.2). Potential outcomes are fixed; the only randomness is assignment with , which may depend on the past (biased-coin and adaptive designs are allowed). With predictable guesses , the augmented IPW pseudo-outcome
is conditionally unbiased for , and Corollary 2 gives for a sub-exponential boundary with scale . This is a design-based, finite-sample, always-valid interval for the running sample ATE, in the spirit of Randomization Inference - Overview. In the paper’s simulation the bound is about twice the CLT width for from to , while the CLT interval fails to cover at many times. Howard et al. note that Optimizely’s two-sample mSPRT relied on asymptotics “which lack rigorous justification”; this construction supplies it.
Running intersection
If is truly constant, is also valid and never wider (equivalently is always valid). But under drift the intersection can become empty; Optimizely’s fix was to reset the experiment. The authors recommend the non-intersected sequence and a time-varying estimand , which stays meaningful under non-stationarity such as novelty effects or day-of-week cycles.
Examples
Width penalty (own computation from Eqs. 2 and 14, , unit variance). Ratio of the confidence-sequence radius to the fixed-sample :
| Normal mixture, | Normal mixture, | Stitched (Eq. 2) | |
|---|---|---|---|
| 1.87 | 4.17 | 2.04 | |
| 1.55 | 1.87 | 2.10 | |
| 1.67 | 1.55 | 2.15 | |
| 1.83 | 1.67 | 2.18 |
The mixture is tightest near and degrades slowly on both sides, so choose about one tenth of the planned sample size. A factor 1.6 in width is a factor of about 2.5 in sample size for the same precision at a fixed time, but the experimenter may stop the moment the interval excludes zero, which for large effects happens far earlier.
import numpy as np
def normal_mixture_cs(x, sigma, alpha=0.05, rho=1000.0):
"""Two-sided anytime-valid CI for the mean of a sigma-sub-Gaussian stream."""
t = np.arange(1, len(x) + 1)
xbar = np.cumsum(x) / t
radius = sigma * np.sqrt((t + rho) * np.log((t + rho) / (alpha**2 * rho))) / t
return xbar - radius, xbar + radius
# Two-arm test with paired arrivals: x = y_treat - y_ctrl, sigma = sqrt(2) * sigma_arm.
# Stop (or ship) the first time the interval excludes 0; validity holds at any stopping time.Reference implementations of all boundaries are in the authors’ confseq R/Python package.
Connections
- The Peeking Problem and Optional Stopping — pointwise intervals fail under monitoring because no boundary survives the LIL.
- Always-Valid p-values and the mSPRT — dual object; the normal mixture boundary equals the Gaussian mSPRT threshold with .
- CUPED and Regression-Adjusted Variance Reduction — predictions in the empirical-Bernstein sequence play the role of the control variate.
- Randomization Inference - Overview and Potential Outcomes Framework — the sequential ATE result is design-based with fixed potential outcomes.
- UCB and Greedy Algorithms for Bandits — UCB indices are one-sided confidence sequences; best-arm identification stops when sequences separate.
- Regret Bounds for Thompson Sampling — TS regret analysis only needs some valid upper confidence bound to exist; time-uniform boundaries are the sharpest generic source of such bounds.
- Multiple Testing Corrections — a Bonferroni split of across metrics keeps simultaneous anytime coverage.