Explore interactive workflow demonstrations, scenario analyses, and example reports that showcase the MMM Framework in action.
Every demo here exercises the same causal measurement loop—pre-specify the model, fit it, validate against real-world experiments, calibrate, and decide—so each page is one rehearsal of the discipline the framework is built around.
Where to spend your time, by who you are and how much of it you have
| You are… | 30 minutes | Half a day |
|---|---|---|
| A skeptical evaluator | The scorecard → The Gauntlet → Aurora 05 | Add the measured backtest, the evaluator's guide, and re-run the quickstart yourself |
| A business sponsor | For Business → the example report | Add Interpreting Results and the calibration loop overview |
| An analyst new to Bayes | Workshop 00 | Workshop 00–03, then the walkthrough |
| A methodologist | Stress 00 → the identification contract | The stress series end-to-end, then the math companion |
Most demos here run on synthetic worlds — deliberately. Real data cannot grade an MMM's attribution, because the true causal effect is unobservable; only data with a known answer key can measure the gap between "all diagnostics green" and "the answer is right," which is precisely what the pressure-testing program does (including publishing the failures). The objection "of course it works on data you generated" is answered two ways: the stress worlds are built to break the model, not flatter it, and the framework is also graded on real data — the Lydia Pinkham pressure test against 50 years of published econometrics.
Step-by-step interactive walkthroughs of the Bayesian modeling process
The philosophy: why honest, transparent iteration beats specification shopping, and the principles that keep a model trustworthy under pressure.
The guided read for first-time analysts: a conceptual tour of why pre-specified, transparent modeling produces more trustworthy results than traditional approaches.
The interactive step-by-step walkthrough: the full Bayesian modeling workflow from question formulation through actionable insights, with live charts and story-to-math connections.
A deliberately messy client CSV — mixed dates, a TOTAL row, missing spend, a negative refund — taken through cleaning, MFF assembly, the EDA gate, fitting, and the validation battery that substitutes for ground truth when none exists. Measured: MASE 0.25 out of time, −6% vs the disclosed answer key.
Deep-dive workflows addressing specific business questions
Measure the true incremental impact of each media channel with honest uncertainty. Includes ROI estimation, cross-channel comparison, and attribution analysis.
Optimize your media budget allocation using saturation curves and uncertainty-aware optimization. Explore what-if scenarios and diminishing returns.
Use model estimates to forecast outcomes under different spending scenarios. Includes uncertainty propagation for honest projections.
The full decision-theoretic stack behind the closed-loop calibration framework, with eight live calculators: utility & risk, EIG, EVOI, the priority matrix, constraint-aware portfolio selection, and a multi-cycle simulator that compounds learning over time.
One story, end to end: Aurora Coffee Co. decides next quarter's budget — and a causal MMM
keeps a dashboard from getting it expensively wrong. Mirrors nbs/00–05.
The dashboard crowns Search (correlation 0.85). The sealed answer key says its true ROAS is 0.66 — a demand-chasing mirage — while the "weak" channels are the brand engines.
The causal DAG, identification, bad-control detection — and why observational adjustment alone is insufficient until a geo-lift experiment pulls Search back to truth.
What is each channel worth? Contributions and ROAS with credible intervals, marginal ROAS, what-if scenarios — and an honest account of what the base model still gets wrong.
What is TV really doing? NestedMMM shows TV is ~fully mediated through awareness; MultivariateMMM weighs the Cold Brew cannibalization claim.
From posterior to boardroom: the report generator's API, sections, and themes — and the doctrine of what never gets stripped out when someone asks for "just the number".
The payoff: the experiment-anchored causal plan beats the dashboard plan by ≈$11.9M/yr on the same budget — a defensible reallocation, with uncertainty.
The feature catalog: role typing, identification analysis, sensitivity & refutation diagnostics, pre-specification locks, and experiment calibration (prior and likelihood routes).
The deep causal-inference treatment: one brand (Veranda Home), many hidden worlds, every claim
graded against a sealed answer key — confounding, structural learning, experiments, calibration,
and the closed measurement loop. Mirrors nbs/causal/causal_00–10.
Naive reads double Social's value; a real Bayesian MMM stays fooled too (+43% on Search). Confounding is a property of the world — and the dashboard can't tell you which world you're in.
Three hidden histories: a measured confounder halves the chasers' bias, a noisy proxy leaves the door ajar, and budget pacing has no door to close at all.
Estimands (marginal graded too), the bad-control trap sprung on purpose, a refutation battery — and the Radio/Print ridge, the model's honest "I don't know".
The brand funnel learned from surveys: binomial awareness tracker, Likert consideration, latent demand — all five structural parameters recovered inside their 90% intervals.
The economy confounds spend and sales, measured only by four noisy indicators. Three rungs of adjustment, honestly graded — including what the joint latent-factor model really buys.
Matched randomized pairs, A/A false-positive rates (one "sophisticated" estimator lies 26% of the time), injected-truth power, and the methodology leaderboard.
A prior is a suggestion; a likelihood is a commitment. One honest experiment repairs a confounded posterior — and the wrong estimand silently voids the test.
Evidence composes: the ridge snaps, an off-panel test disciplines the curve, and a confident wrong certificate is caught by a 2.75σ tension check — then merged anyway, to price the damage.
Model-anchored power, near-free budget-neutral learning, the Pareto frontier — and the design engine's most valuable answer: "don't run this test."
EIG × EVOI quadrants, the EVPI ceiling on the learning budget, evidence that decays on a half-life, and a twelve-month calendar with a validation track.
Three cycles of fit → prioritize → design → measure → calibrate: error falls, EVPI shrinks, and the budget decision goes from value-destroying to near-optimal.
Bayesian MMM from zero: no Bayesian background assumed, every term glossed in plain English,
every estimate graded against a known answer key. Mirrors nbs/workshop_00–05.
Probability as belief, Bayes' rule on a grid, credible intervals, and your first derived quantity: "is the new creative better?" answered as P(B > A) and an uplift distribution.
The MMM prior zoo — and the doctrine that a prior on a parameter is really a prior on behavior, shown as slider-driven fans of implied adstock and response curves.
Drive a Metropolis sampler in your browser, break it with the step-size slider, and watch R-hat and ESS — computed live on your own chains — catch every failure you caused.
MMM vocabulary, the config built choice by choice, the diagnostics gate — and a first fit that recovers the known answer at ~7% median error, checked by you.
The posterior as a table of internally-consistent worlds: HDIs, forest plots, the β–λ trade-off at r ≈ −0.8 that proves derived quantities are computed per draw.
Compute per draw, summarize last. ROAS as a distribution, marginal vs average rankings that genuinely disagree, and a reallocation that wins with P = 0.98.
The mathematics under the hood — every transform, prior, and likelihood term derived and
graphed, grounded in the framework's real code. Mirrors nbs/math_00–06.
The full additive μt, the Normal likelihood, the complete default-prior table, and a live draw from the prior — the index for the set.
Carryover as convolution: geometric IIR vs normalized FIR kernels, the 1/(1−α) multiplier, and what the legacy Beta(2,2) two-α blend (the pre-2026 default) can and cannot express.
The core exponential curve (strictly concave — not S-shaped), half-saturation at ln2/λ, the Hill alternative with real thresholds, and marginal returns.
Fourier order as a bias–variance dial, the corrected Prophet-style piecewise trend (slope changes, not level jumps), and B-splines as a partition of unity.
Priors → prior predictive → NUTS → convergence → posterior intervals → PPC → additive decomposition — the full inference pipeline run on Aurora.
The equifinality trap, the prior route (design factor + Gamma moment-matching) and the likelihood route — and the demo that pulls Search's ROAS from 2.8 to its true 0.66.
The structural equations behind NestedMMM mediation, MultivariateMMM cross-effects (ψ in the mean vs correlation in the noise), and CombinedMMM as their synthesis.
The model-free sequential loop — planning a budget from designed geo experiments alone, one wave at a time. Two baked notebooks behind the continuous-learning guide.
The Nomi walkthrough: why the dashboard lied, one designed wave, the funding line and synergy map, reallocation with confidence, the stop decision, and the re-test clock — the full cycle a media analyst runs.
Recovery against known truth, the central-composite design, Thompson allocation, Laplace-KG vs pure-EIG acquisition, ENBS stopping — plus the misspecification study behind the “trust the ranking, not the magnitudes” rule.
The framework's known strengths and weaknesses, measured: synthetic worlds engineered to
break the model, graded against known causal truth. Mirrors nbs/stress_00–06.
16 worlds, 8 silent failures — every one with clean convergence diagnostics. The interactive scorecard, the strengths/weaknesses ledger, and the fix ladder.
The full analyst workflow — EDA pre-flight, pre-specified model, diagnostics gate, calibration — on a world that is realistic but kind. The gentle companion to the Gauntlet.
A positive control recovers at 7% error — and an equally-green confounded fit is badly wrong on Search. What each diagnostic can and cannot see.
Delayed carryover vs a kernel that can't bend (median error 88%); Hill thresholds vs a concave default (+40%, all green); 15× spend spikes (coverage 0%).
A trend break plus a reactive media ramp sends TV to −63%; evolving seasonality scrambles the channel split while totals stay right.
Demand-chasing spend (+110% on Search, all diagnostics green), the noisy-proxy trap, the collinear seesaw — and the one lift test that closes the gap.
NestedMMM recovers what the base model loses — and over-credits its shared-mediator sibling 2.6×; a cross-effect flips sign with the prior; a wrong mediator fits perfectly.
Everything at once: ranking fully inverted, Search 8× overstated — and the full recovery workflow, each step measured, ending at Search ROAS 5.59 → 0.65 (truth 0.66).
National numbers green while per-geo errors hit +330% and the regional ROI ranking fully inverts (ρ = −1) — plus the cross-geo adstock bug this series found and fixed.
Forecast accuracy, runtime, and a real-data pressure test — measured in baked, seeded notebooks whose recorded artifacts the docs quote.
Rolling-origin backtest: 52 out-of-time weekly forecasts, MAPE 3.0% vs 13.8% seasonal-naive (MASE 0.24), honest interval coverage — plus the trend-break negative control where the harness correctly reports failure.
Wall-clock by data shape, samplers, draws, geo panels, and extension models on disclosed hardware: a production-size national fit in ~15 s, NumPyro ~3× vs sequential PyMC — every docs runtime claim traced to this artifact.
The framework on the most-studied real advertising series in econometrics, graded against 50 years of published estimates: carryover sides with Clarke/Hanssens over Palda, ROAS inside the literature bracket, and a forecast grade reported honestly.
Printable engagement documents, generated by the reporting module
See what a complete MMM Framework report looks like
Pre-generated report examples showing different themes and formats produced by the framework
Install the framework and start creating models with honest uncertainty quantification.