A marketing mix model answers a causal question β "what would sales have been without this media?" β and an answer you can defend takes discipline, not just software. This framework enforces four: the model design is pre-registered before results are seen, the methodology is pressure-tested in public against known ground truth, the model's stated uncertainty is verified, not assumed, and the estimates are calibrated against real experiments in a closed loop.
Each path is an ordered route, not a single page. Want it mapped by time budget? See the Recommended Paths matrix.
What MMM can and cannot promise, the seven questions to ask any vendor, and why a tight number is not a true number.
Route: For Business › Example Report › Interpreting Results
Install, fit your first model in seconds, then take the workshop series from regression to posterior.
Route: Getting Started › Workshop series › Aurora Tour
16 synthetic worlds with known causal truth, 8 published silent failures, and the identification contract behind the claims.
Route: Pressure Testing › Stress series › Identification › Math series
License, version, maturity tiers, support reality, measured runtime and accuracy β the procurement facts on one page.
Route: Evaluator › The Gauntlet › Measured backtest
Fifteen essays on modern measurement β causal inference, geo experiments, Bayesian experimental design β written from the primary literature.
Route: Research index › the essays, in order
When you run many models and only report the "good" ones, you're painting targets around arrows.
Lots of things move together: ice cream sales and sunburns rise in the same weeks without one causing the other, and holiday demand lifts both ad spend and sales at once. Separating coincidence from contribution takes discipline β and traditional MMM often skips it, running dozens of model variations and reporting the one that "makes sense."
Watch the dartboard: each throw is a model specification. Only the bullseyes get reported; the misses are quietly discarded, so the reported accuracy is an illusion.
This framework takes the opposite path. Variable roles are declared up front, the model design is locked in a pre-registered analysis plan before results are seen, and answers are checked against real-world experiments such as regional holdout tests. The result is an estimate of incremental impact β what your media actually caused. Read the full argument, or see what specification shopping costs the business.
The darts that miss are models with "unrealistic" resultsβquietly discarded.
π―
Every vendor's deck says "Bayesian" and "calibrated." These four disciplines make the words mean something β each one enforced by the workflow, and each one measured in public rather than asserted.
Specification shopping β running many models and reporting the flattering one β invalidates the very uncertainty statements that make measurement useful. The antidote is a commitment device: variable roles, priors, functional forms, and decision criteria are written into an analysis plan before the model sees the data's verdict, so there is no quiet path to a more flattering model later. Experiments get the same treatment β design, estimator, and primary outcome locked in a signed pre-registration memo before launch, then tracked through a lifecycle from draft to calibrated. The platform can even generate a pre-fit design readout recording every prior, assumption, and prior-predictive check before the final fit runs.
Convergence diagnostics validate the computation, not the causal claim β a model can be confidently, quietly wrong. So the methodology is attacked before it is trusted: 16 synthetic markets with known causal ground truth, each breaking one real-world assumption, graded and published β including the eight worlds where attribution failed silently while every standard check stayed green. A four-test causal refutation suite ("a tripwire, not a verdict"), prior and posterior predictive checks, and a prior-dominated-posterior diagnostic complete the critique toolkit β each documented with what it can and cannot see.
A model is calibrated when its stated uncertainty is accurate β when 90% intervals contain the truth about 90% of the time. Wide intervals are honest communication, but only if the width itself can be trusted. The framework verifies it: simulation-based calibration checks that the posterior's uncertainty statements are themselves honest, posterior-predictive coverage is graded against nominal, and rolling-origin backtests confront the intervals with data the model never saw. Approximate fits (MAP, ADVI, Pathfinder) are labeled what they are β not calibrated, re-fit with NUTS before deciding β so a quick check is never mistaken for a defensible posterior.
Only evidence from outside the observational data can catch a silent failure. The framework prices which experiment buys the most learning β expected information gain in bits, expected value of information in dollars β then designs pre-registered geo-lift, matched-market, and budget-neutral flighting tests directly from the posterior. Readouts fold back in β as informed priors on a channel's coefficient, or directly in the likelihood, where they update the coefficient, the saturation curve, and the adstock kernel jointly. Calibration is surgical, not contagious: it corrects what was tested, and untested channels keep their bias β which is why the loop keeps running.
The four disciplines close into one loop: a pre-registered model is critiqued, its uncertainty is verified, and its claims are tested against reality β then the experiment results become the next cycle's priors. Every measured number above traces to a seeded, re-runnable notebook: measured, not asserted.
Prior beliefs combine with data to produce a posterior β and the width of that posterior is information, not weakness. Adjust the levers and see what the data can and cannot pin down.
One model fit is a snapshot. The framework's operating rhythm is a cycle that keeps sharpening your answers β like a research budget for your media plan, spent exactly where learning changes the next decision most.
Estimate incremental impact for every channel, with honest uncertainty ranges instead of single guaranteed numbers.
Pinpoint where uncertainty costs the most, in dollars β expected information gain and expected value of information rank what to learn next.
Pre-registered tests that buy the most learning: regional holdout (geo lift), matched-market, or budget-neutral flighting designs β success criteria locked before launch.
Experiment readouts flow into a calibrated refit β the model proposes, a real-world test disposes, and the estimates tighten.
Shift spend with the sharper answer, accounting for carryover, saturation, and genuine remaining uncertainty.
Information decays as markets shift, so the cycle flags when a past answer needs re-testing β and the loop repeats.
Every test ships with a pre-registration memo β estimand, design, power, stopping rule β signed before launch. See how the loop works in the calibration loop guide, or follow a worked calibration decision end to end. No usable history to fit a model on yet? The continuous-learning loop runs the same rhythm model-free, straight from designed geo experiments.
The measurement loop runs in a modern web application β and everything it does is also available as a Python library for teams who prefer code.
Your home base: where you are in the measurement cycle, headline KPIs, next-best actions, and a coverage map of what has been validated.
A priority matrix of what to test next, a lifecycle board from draft to calibrated β every test pre-registered before results are seen β and a studio for designing the tests themselves.
Cycle-over-cycle trajectories, estimand-by-estimand model comparisons, a log of where the model and experiments agree, and a timeline of every model run.
Turn the latest fit into a plan: optimal allocation with per-channel constraints, geo splits, a forward flighting calendar, and what-if scenarios.
No usable history? Learning programs run waves of designed geo experiments and steer spend model-free β with a funding line, synergy map, and a stopping rule.
A chat-based analyst assistant that validates data, fits models, and runs a one-click validation battery β convergence, priorβposterior learning, posterior-predictive checks, and simulation-based calibration β before drafting client-ready reports.
Take the platform tour, or start with the Python library.
Meridian, Robyn, and PyMC-Marketing are serious tools built by serious teams. Here is the difference in approach β stated plainly, the way we would want a vendor to state it to us.
Robyn searches thousands of model candidates with an evolutionary optimizer and asks the analyst to choose from a Pareto front of finalists. That selection step is exactly where specification shopping lives β and exactly what this framework locks down. Here there is one pre-registered Bayesian model, and uncertainty comes from its full posterior, not from which finalist got picked.
Meridian is a capable Bayesian geo-MMM, and its lift-test calibration is real. This framework operates the rest of the loop: it prices which experiment to run next (expected information gain, expected value of information in dollars), pre-registers the design, tracks its lifecycle, and folds the readout back in. It also models mediation β TV driving search driving sales β and multiple KPIs jointly, structures most MMM stacks do not represent.
This framework is a standalone PyMC 6 engine β it does not subclass or depend on PyMC-Marketing (it can optionally interoperate with a PyMC-Marketing model for reporting). PyMC-Marketing provides excellent Bayesian MMM foundations; we built a separate engine and added what a production measurement program needs around it: causal guardrails (declared variable roles, refutation checks), the experiment loop, a full web platform, and an AI analyst workspace.
Every vendor's deck says "Bayesian" and "calibrated." Two claims here are harder to make and easy to verify: the model design is locked before results are seen, and the methodology is pressure-tested in public against synthetic markets with known ground truth β including the worlds where it struggles. See the pressure-test scorecard.
Everything you need to go from your first model fit to a quarterly closed-loop measurement program.
Install the framework, fit your first Bayesian MMM, and walk through a complete code example.
Step-by-step guidance for statistically sound MMMs β the pre-registered analysis plan, priors, hierarchy, diagnostics, and honest iteration.
How to read MMM outputs, communicate uncertainty, and translate posteriors into confident budget decisions.
The measured scorecard: 16 synthetic worlds with known causal truth, 8 published silent failures, and the fix ladder of recovery moves β each measured before and after.
The seven causal assumptions behind any MMM β stated formally, labeled testable or untestable, and priced against the stress-test scorecard.
All models are wrong, some are useful: honest iteration vs. specification shopping, when to stop, and how to communicate what remains uncertain.
The disciplined process: priors, prior predictive checks, sampling diagnostics, posterior predictive checks, and interval calibration (PIT).
Rolling-origin out-of-time accuracy, measured not asserted β including the negative control where the harness honestly fails.
What replaces the truth column on client data: priorβposterior learning, out-of-time backtests, stability checks, and interval honesty.
Wire experiments and the MMM into one self-correcting cycle β EIG, EVOI, three calibration mechanisms, and the virtuous loop.
An interactive workshop: live calculators for the Bayesian update, EIG, EVOI in dollars, experiment portfolios, and geo-test power.
Learn channel response and synergies straight from designed geo experiments β no fitted MMM required, with a built-in stopping rule.
DAGs, do-calculus, and counterfactuals β the causal scaffolding behind every defensible MMM specification.
When to include precision controls, why confounders are non-negotiable, and how shopping for variables backfires.
The risks of specification shopping, the seven questions to ask any modeling partner, and what defensible measurement looks like.
Interactive workflow demonstrations, scenario analyses, and an example MMM report from a Q4 2025 fit.
Author bespoke Bayesian models in-app, prove them against a compatibility suite, and publish versions any analyst can fit through the agent.
Math and implementation details for every model in the framework β standard, nested, multivariate, and combined.
Module-by-module API docs (Sphinx). Browse the source on GitHub or build locally with make html in docs/api.
Definitions of every term used across the docs β adstock, EIG, EVOI, MFF, ITT, MDE, ROPE, and more.
Common questions about MMMs, Bayesian methods, and the trade-offs the framework makes deliberately.
Current version, API stability tier per module, versioning policy, and release notes.
Background on the framework, its design principles, and the audiences it serves.
The framework is open source, the methodology is public, and every measured claim traces to a notebook you can re-run. Start with the documentation or go straight to the evidence.