Evaluating and Comparing

Routing Summary

Chapters 8–10 of Gelman, Vehtari & McElreath (2026): checking a fitted model, measuring sensitivity, comparing models by predictive performance, expanding rather than selecting, and the scientific context that makes iterated analysis defensible. 16 notes.

Concept Map

ConceptNoteTypeDepends OnKey Result
Visualizing a high-dimensional posteriorVisualizing High-Dimensional Inferenceconcept—Marginals mislead under collinearity; look at joints and derived quantities
Posterior predictive checkingPosterior Predictive CheckingconceptPrior Predictive CheckingCompare to graphically, not by -value
Cross validationCross Validation CheckingtheoremPosterior Predictive CheckingPSIS-LOO with Pareto ; flags failure
Pointwise influenceInfluence of Individual Data PointsconceptCross Validation CheckingFigure 8.10; which points move the fit
Prior/likelihood sensitivityInfluence of Likelihood and PriorconceptCross Validation CheckingPower-scaling with priorsense; Figures 8.11–8.13
Big data need big modelsBig Data Need Big Modelsconcept—More data exposes more model misfit, not less
Model topologyTopology of ModelsdefinitionBig Data Need Big ModelsModels exist in relation to each other
Visual model comparisonComparing Models VisuallyconceptTopology of ModelsFigures 9.1, 9.2
Predictive model comparisonModel Selection Using Predictive PerformancetheoremCross Validation CheckingEq. 9.1–9.5; elpd differences and their standard errors
Selection overfittingModel Selection and OverfittingconceptModel Selection Using Predictive PerformanceThe search itself overfits noisy CV estimates
StackingStacking and Predictive Model AveragingdefinitionModel Selection Using Predictive PerformanceWeight by predictive performance, not posterior probability
Predictive consistencyModel Expansion - Predictive Consistency and CoherencedefinitionTopology of ModelsExpansion should not inflate the prior predictive
Statistical vs. scientific inferenceStatistical and Scientific Inferenceconcept—Figures 10.1, 10.2
ToolingSoftware Assisted WorkflowconceptStatistical and Scientific InferenceWhat automation should and should not do
ReplicationThe Replication Crisis and Multiple Levels of VariationconceptStatistical and Scientific InferenceVariation exists at several levels; most designs measure one
Virtual replicationSimulated-Data Experimentation as Virtual ReplicationconceptThe Replication Crisis and Multiple Levels of VariationFigures 10.3–10.6; the answer to forking-paths objections

Notes

Sources

  • Gelman Vehtari McElreath 2026 - Bayesian Workflow (book) — Chapters 8–10, pp. 137–190

See Also