From Inference to Decision

Summary

Why the significance threshold is the wrong decision rule, and what to use instead. The arithmetic that anchors the argument: an estimate exactly one standard error from zero — “only weak evidence” by conventional standards — still implies an 84% posterior probability that the effect is positive (76% under a moderately informative prior). The chapter reanalyzes the ORBITA heart-stent trial, published as a null finding at , and shows that under either a flat prior or a Cochrane-derived meta-analytic prior the data support a small positive effect — probability 0.71 or 0.80 of an effect between 0 and 30 seconds.

Overview

The canonical setup (Ch. 7.3, pp. 130-131)

An unbiased estimate with standard error : . With a uniform prior, — “conveying that we are treating the estimate and standard error as sufficient statistics.”

With prior — “you have no reason to believe the effect is positive or negative, but small effects are more likely than large effects, and represents the scale of effects that might be expected”:

Two calibrations of offered:

  • A/B test of a small marketing innovation, outcome = log spending, mature industry: , “implying that we would not expect any intervention to increase or decrease average spending by much more than 1%.”
  • Cancer treatment, typical tumor size ~100 grams: , “implying that it is unlikely but possible that the treatment could cause tumor sizes to double on average or to bring them to zero.”

“Our point here is not that these priors are ‘correct’ but that assumptions about the population of possible effect sizes are relevant to decision making, so we should be thoughtful and open in our assumptions about this population.”

The one-standard-error calculation

Suppose and — exactly one standard error from zero.

“It is fair to say this is only weak evidence in favor of a positive effect: it would be no surprise to see an estimate this large even if the true were zero. On the other hand, under this model the posterior probability is 84% that is positive.”

With a moderately informative , the posterior probability that is

“Including this prior decreases the posterior odds of a positive effect from 5:1 to 3:1.”

The point is not that 84% is high — it is that the significance framework maps this to “no effect,” discarding the 5:1 odds entirely.

Main Content

The Bayesian answer, and three reasons to stop short of it

Propagating the posterior into a decision (Ch. 7.3, p. 131)

“From the Bayesian perspective, the answer is clear: take the posterior distribution and propagate it as necessary to any decisions.”

The A/B example. Treatment A is the status quo, B a marketing idea, the relative dollars gained per potential customer under B. The idea applies to potential customers and costs to implement. Go with B if the expected net gain is positive:

The three assumptions this rule requires, stated explicitly:

  1. Stakes are small enough for expected monetary value. “If your company is making many small independent decisions, none of which has an extremely variable outcome, the total effect on the bottom line will be approximately the sum of their expected values, from the law of large numbers.”
  2. No other known or predictable outcomes — “as would arise, for example, if acting on this marketing idea would preclude the development or implementation of other, potentially better, innovations; or, conversely, if implementing the idea would open the door to further developments.” So “is intended to represent the net cost of implementation, including expected opportunity costs or benefits.”
  3. The model is an approximation. “Even the discussion of a ‘treatment effect’ is a simplification, as real effects will depend on situations which change over time.”

Three legitimate reasons to separate inference from decision (Ch. 7.3, pp. 131-132)

“It is common practice to use the statistical information obtained about to make an inferential summary without reference to the ultimate decision problem.”

  1. Costs and benefits might be unknown — “rather than propagating uncertainty all the way to the end of the process it makes sense to pause and just say what is known about .”
  2. Outcomes can be multidimensional — “as when considering a new medical treatment that is expensive but could save lives. Ultimately we do think it can make sense to put dollars and lives on a common scale (Lin et al. 1999), but we will generally be interested in the estimated costs and benefits, not just in the decision recommendation.”
  3. Different people do different jobs — “the experimenters, data analysts, and decision makers could be three different groups of people, in which case the job of the statistician is to aid in design and analysis and then to stop and supply the necessary information for others to make the decision.”

The third reason is “related to the idea of multiple imputation, in which missing data are imputed without direct reference to the purpose for which the completed datasets will be used” (Meng 1994; Rubin 1996). See Missing Data Models.

Against classical thresholding

Two well-known problems with significance-based decisions (Ch. 7.3, p. 132)

Classical thresholding: if , act as if the point estimate is the true effect; otherwise act as if it is zero.

  1. “Selection on statistical significance leads effect sizes to be overestimated” — type M (magnitude) errors (Gelman and Carlin 2014).
  2. “It throws away information to ignore results that do not cross the significance threshold” — “as discussed above, even an estimate that is only 1 standard error from zero, and thus easily explainable by chance, can still correspond to a high probability of identifying the sign of the effect.”

“Using the posterior mean as a point estimate avoids these problems.”

But the naive Bayesian threshold is too strict

“A natural Bayesian counterpart to classical thresholding would be to declare ‘statistical significance’ or ‘high confidence’ if the posterior mean is more than 2 posterior standard deviations away from zero.

The problem here is that, to the extent you believe your posterior distribution, this threshold is very strict, corresponding to at least a 97.5% posterior probability of getting the correct sign. Even if , the posterior odds are still 5:1 of identifying the sign of correctly.

Decades of classical estimates have trained us to distrust point estimates, but such distrust is not necessarily appropriate for inferences that have already been partially pooled.”

This is subtle and easy to miss: the reason classical point estimates deserve distrust is that they are unregularized. A partially pooled posterior mean has already had the exaggeration removed — see The stakes: exaggeration factors.

The recommendation: present everything (Figure 7.9)

“In general we recommend presenting all estimates of potential interest, along with uncertainties, rather than selecting based on a cutoff of any kind. Bayesian inferences account for uncertainty, and let’s take advantage of that.”

Figure 7.9 shows two designs for displaying many posterior estimates and uncertainties on a single plot (from Gelman and Margalit 2021, and Mitchell, Gelman, Ross, et al. 2018).

“Larger studies could require a more elaborate series of graphical displays, but we believe this would be worth the effort, as compared to the usual practice of partially reporting results based on thresholds.”

Examples

ORBITA — reanalyzing a "null" heart-stent trial (Al-Lamee et al. 2017; Ch. 7.3, pp. 132-134)

The published result. An estimated increase in treadmill time of 16.6 seconds, se 12.7, — “not ‘statistically significant’ and was thus reported as a null finding, to the extent that people expressed surprise that the treatment was ineffective. For example, a news report characterized the result as ‘unbelievable … stunned leading cardiologists by countering decades of clinical experience.‘”

The improved analysis. “The statistical analysis in the published paper was statistically inefficient.” After correction (Gelman, Carlin, and Nallamothu 2019): estimate 21.3, se 12.6, .

The even-handed reading of the -value:

  • On one hand, “the significance level does imply that an estimate this large would not be a surprise even in the absence of any effect; that is, the magnitude of the estimate, relative to its uncertainty, is compatible with there being no effect.”
  • On the other hand, “the estimated effect is positive. The data are more compatible with a positive effect than a negative effect; moderately large positive effects are compatible with the data, whereas any compatible negative effects would be very small.”

Two reference points for judging the magnitude — both essential to the interpretation:

  1. Against the outcome’s own spread. Mean pre-treatment treadmill time was 506 seconds, sd 188. So 21.3 is “approximately one-tenth of a standard deviation, which is not a huge shift in the distribution but can be meaningful for a treatment that represents one component of a larger regimen and is not intended singlehandedly to resolve a medical condition.”
  2. Against the design. “The study was designed to have sufficient power to detect an effect of 30 seconds, which is a bit less than the benefits estimated from single anti-anginal agents.” This gives a natural three-way partition of the parameter space.

Analysis 1 — flat prior (reproducing the classical analysis):

RegionPosterior probability
Effect negative5%
Effect between 0 and 30 s71%
Effect greater than 30 s24%

“Using this default analysis, it seems fair to conclude that the experiment shows a positive effect for stents in this setting, although the effect is most likely small and lower than was hoped.”

Analysis 2 — the Cochrane meta-analytic prior (van Zwet and Gelman 2022; see Constructing Priors for Effect Sizes):

“If we assume that the standard error provides no additional information about the effect size, we can combine the prior with the normal likelihood to yield a normal-mixture posterior, although the weights will no longer be 0.42 and 0.58. Rather than working out the algebra, we simply combine the distributions in Stan and simulate.”

RegionPosterior probability
Effect negative0.07
Effect small (0-30 s)0.80
Effect moderately large (>30 s)0.13

The conclusion: “So if considered as a sample from the experiments in the Cochrane database, the results again assign the highest probability to a small positive effect. We think this would be a better summary than the statement in the published paper that the treatment ‘did not increase exercise time by more than the effect of a placebo procedure’” — consistent with Kass (2011) on interpreting statistical inferences pragmatically in light of their assumptions.

Note how little the two priors differ in conclusion. Both say “small positive effect, most likely.” The informative prior mainly shifts mass from “>30 s” to “0-30 s” — it regularizes the optimistic tail, not the sign.

Different perspectives on modeling and prediction (§7.4)

Three inferential goals (Ch. 7.4, p. 134)

PerspectiveGoalWhen computation can stop
Traditional statisticalAccurately summarize the posterior for a model chosen ahead of time”as long as necessary to reach approximate convergence”
Machine learningPrediction, not parameter estimation”when cross validation prediction accuracy has plateaued”
Model explorationTrying out a series of models, “many of which will have terrible fit to data, poor predictive performance, and slow convergence”approximations are attractive — but see the caveat

What each implies:

  • Traditional: “run computation for a long time, using approximations only when absolutely necessary. Another way of saying this is that in traditional statistics, the approximation might be in the choice of model rather than in the computation.”
  • Machine learning: “pick an algorithm that trades off predictive accuracy, generalizability, and scalability, so as to make use of as much data as possible within a fixed computational budget.”
  • Model exploration: “we want to cycle through many models, which makes approximations attractive. But there is a caveat: if we are to efficiently and accurately explore the model space rather than the algorithm space, we require any approximation to be sufficiently faithful as to reproduce the salient features of the posterior.”

What the distinction is not: “The distinction here is not about inference vs. prediction, or exploratory vs. confirmatory analysis. Indeed all parameters in inference can be viewed as some quantities to predict, and all of our modeling can be viewed as having exploratory goals (Gelman 2003). Rather, the distinction is how much we trust a given model and allow the computation to approximate.”

Why problems seem obvious only in hindsight

“As we illustrate with the case studies in Chapters 25 and 30, problems with a statistical model often seem obvious in hindsight, but we need the workflow to identify them and understand the obviousness.

Another important feature of these examples … is that particular challenges in modeling arise in the context of the data at hand: had the data been different, we might never have encountered these particular issues, but others might well have arisen. This is one reason that subfields of applied statistics advance from application to application, as new wrinkles become apparent in existing models.”

Connections

See Also