🎯 Open Source · Apache-2.0

Honest, Causal Marketing Measurement

A marketing mix model answers a causal question β€” "what would sales have been without this media?" β€” and an answer you can defend takes discipline, not just software. This framework enforces four: the model design is pre-registered before results are seen, the methodology is pressure-tested in public against known ground truth, the model's stated uncertainty is verified, not assumed, and the estimates are calibrated against real experiments in a closed loop.

Choose your path

Each path is an ordered route, not a single page. Want it mapped by time budget? See the Recommended Paths matrix.

The Hidden Problem with Traditional MMM

When you run many models and only report the "good" ones, you're painting targets around arrows.

Coincidence Is Not Contribution

Lots of things move together: ice cream sales and sunburns rise in the same weeks without one causing the other, and holiday demand lifts both ad spend and sales at once. Separating coincidence from contribution takes discipline β€” and traditional MMM often skips it, running dozens of model variations and reporting the one that "makes sense."

Watch the dartboard: each throw is a model specification. Only the bullseyes get reported; the misses are quietly discarded, so the reported accuracy is an illusion.

This framework takes the opposite path. Variable roles are declared up front, the model design is locked in a pre-registered analysis plan before results are seen, and answers are checked against real-world experiments such as regional holdout tests. The result is an estimate of incremental impact β€” what your media actually caused. Read the full argument, or see what specification shopping costs the business.

Models Run 0
Models Reported 0

This Is Specification Shopping

The darts that miss are models with "unrealistic" resultsβ€”quietly discarded.

Four Disciplines of Honest Causal MMM

Every vendor's deck says "Bayesian" and "calibrated." These four disciplines make the words mean something β€” each one enforced by the workflow, and each one measured in public rather than asserted.

Discipline 01 · Pre-Registration

The Design Is Locked Before Results Are Seen

Specification shopping β€” running many models and reporting the flattering one β€” invalidates the very uncertainty statements that make measurement useful. The antidote is a commitment device: variable roles, priors, functional forms, and decision criteria are written into an analysis plan before the model sees the data's verdict, so there is no quiet path to a more flattering model later. Experiments get the same treatment β€” design, estimator, and primary outcome locked in a signed pre-registration memo before launch, then tracked through a lifecycle from draft to calibrated. The platform can even generate a pre-fit design readout recording every prior, assumption, and prior-predictive check before the final fit runs.

Simulated Run 20 specifications and report only the best: for a true effect of 0.05, the expected reported estimate is 0.35 β€” a 7× overestimate. Pre-registration is the lock.
Discipline 02 · Model Critique

Attack the Model Before You Trust It

Convergence diagnostics validate the computation, not the causal claim β€” a model can be confidently, quietly wrong. So the methodology is attacked before it is trusted: 16 synthetic markets with known causal ground truth, each breaking one real-world assumption, graded and published β€” including the eight worlds where attribution failed silently while every standard check stayed green. A four-test causal refutation suite ("a tripwire, not a verdict"), prior and posterior predictive checks, and a prior-dominated-posterior diagnostic complete the critique toolkit β€” each documented with what it can and cannot see.

Measured 8 of 16 stress worlds failed silently β€” materially wrong attribution with R-hat ≤ 1.02 and zero divergences on every one.
Discipline 03 · Model Calibration

Uncertainty You Can Quote

A model is calibrated when its stated uncertainty is accurate β€” when 90% intervals contain the truth about 90% of the time. Wide intervals are honest communication, but only if the width itself can be trusted. The framework verifies it: simulation-based calibration checks that the posterior's uncertainty statements are themselves honest, posterior-predictive coverage is graded against nominal, and rolling-origin backtests confront the intervals with data the model never saw. Approximate fits (MAP, ADVI, Pathfinder) are labeled what they are β€” not calibrated, re-fit with NUTS before deciding β€” so a quick check is never mistaken for a defensible posterior.

Measured Out-of-time backtest: 3.0% MAPE, with 80% intervals covering 90% of held-out weeks β€” and on the trend-break negative control the harness honestly fails (58% coverage). A backtest that cannot fail cannot validate.
Discipline 04 · Experimental Calibration

The Model Proposes, an Experiment Disposes

Only evidence from outside the observational data can catch a silent failure. The framework prices which experiment buys the most learning β€” expected information gain in bits, expected value of information in dollars β€” then designs pre-registered geo-lift, matched-market, and budget-neutral flighting tests directly from the posterior. Readouts fold back in β€” as informed priors on a channel's coefficient, or directly in the likelihood, where they update the coefficient, the saturation curve, and the adstock kernel jointly. Calibration is surgical, not contagious: it corrects what was tested, and untested channels keep their bias β€” which is why the loop keeps running.

Measured In the gauntlet world, one lift test pulls Search ROAS from 5.59 to 0.65 against a ground truth of 0.66 β€” while an uncalibrated sibling channel stays wrong.

The four disciplines close into one loop: a pre-registered model is critiqued, its uncertainty is verified, and its claims are tested against reality β€” then the experiment results become the next cycle's priors. Every measured number above traces to a seeded, re-runnable notebook: measured, not asserted.

Watch Honest Uncertainty Update

Prior beliefs combine with data to produce a posterior β€” and the width of that posterior is information, not weakness. Adjust the levers and see what the data can and cannot pin down.

Adjust Parameters

Value: 0.3
Value: 0.5
Value: 50
Value: 0.5

Measurement Is a Loop, Not a Report

One model fit is a snapshot. The framework's operating rhythm is a cycle that keeps sharpening your answers β€” like a research budget for your media plan, spent exactly where learning changes the next decision most.

T₀ · Fit

Fit the Model

Estimate incremental impact for every channel, with honest uncertainty ranges instead of single guaranteed numbers.

T₁ · Find

Find the Expensive Unknowns

Pinpoint where uncertainty costs the most, in dollars β€” expected information gain and expected value of information rank what to learn next.

T₂ · Run

Run the Right Experiment

Pre-registered tests that buy the most learning: regional holdout (geo lift), matched-market, or budget-neutral flighting designs β€” success criteria locked before launch.

T₃ · Feed Back

Feed the Result Back

Experiment readouts flow into a calibrated refit β€” the model proposes, a real-world test disposes, and the estimates tighten.

T₄ · Reallocate

Reallocate Budget

Shift spend with the sharper answer, accounting for carryover, saturation, and genuine remaining uncertainty.

T₅ · Re-evaluate

Re-evaluate as Knowledge Ages

Information decays as markets shift, so the cycle flags when a past answer needs re-testing β€” and the loop repeats.

Every test ships with a pre-registration memo β€” estimand, design, power, stopping rule β€” signed before launch. See how the loop works in the calibration loop guide, or follow a worked calibration decision end to end. No usable history to fit a model on yet? The continuous-learning loop runs the same rhythm model-free, straight from designed geo experiments.

The Platform

The measurement loop runs in a modern web application β€” and everything it does is also available as a Python library for teams who prefer code.

Take the platform tour, or start with the Python library.

Honest About the Alternatives

Meridian, Robyn, and PyMC-Marketing are serious tools built by serious teams. Here is the difference in approach β€” stated plainly, the way we would want a vendor to state it to us.

vs. Robyn (Meta)

One Pre-Specified Model, Not a Menu

Robyn searches thousands of model candidates with an evolutionary optimizer and asks the analyst to choose from a Pareto front of finalists. That selection step is exactly where specification shopping lives β€” and exactly what this framework locks down. Here there is one pre-registered Bayesian model, and uncertainty comes from its full posterior, not from which finalist got picked.

vs. Meridian (Google)

Calibration Plus the Planning Loop

Meridian is a capable Bayesian geo-MMM, and its lift-test calibration is real. This framework operates the rest of the loop: it prices which experiment to run next (expected information gain, expected value of information in dollars), pre-registers the design, tracks its lifecycle, and folds the readout back in. It also models mediation β€” TV driving search driving sales β€” and multiple KPIs jointly, structures most MMM stacks do not represent.

vs. PyMC-Marketing

A Separate Engine, Interoperable by Design

This framework is a standalone PyMC 6 engine β€” it does not subclass or depend on PyMC-Marketing (it can optionally interoperate with a PyMC-Marketing model for reporting). PyMC-Marketing provides excellent Bayesian MMM foundations; we built a separate engine and added what a production measurement program needs around it: causal guardrails (declared variable roles, refutation checks), the experiment loop, a full web platform, and an AI analyst workspace.

Every vendor's deck says "Bayesian" and "calibrated." Two claims here are harder to make and easy to verify: the model design is locked before results are seen, and the methodology is pressure-tested in public against synthetic markets with known ground truth β€” including the worlds where it struggles. See the pressure-test scorecard.

Explore the Documentation

Everything you need to go from your first model fit to a quarterly closed-loop measurement program.

Start here

Getting Started

Install the framework, fit your first Bayesian MMM, and walk through a complete code example.

For modelers

Modeling Guide

Step-by-step guidance for statistically sound MMMs β€” the pre-registered analysis plan, priors, hierarchy, diagnostics, and honest iteration.

For media planners & CMOs

Interpreting Results

How to read MMM outputs, communicate uncertainty, and translate posteriors into confident budget decisions.

Model critique

Pressure Testing

The measured scorecard: 16 synthetic worlds with known causal truth, 8 published silent failures, and the fix ladder of recovery moves β€” each measured before and after.

Model critique

Identification Assumptions

The seven causal assumptions behind any MMM β€” stated formally, labeled testable or untestable, and priced against the stress-test scorecard.

Methodology

Scientific Modeling

All models are wrong, some are useful: honest iteration vs. specification shopping, when to stop, and how to communicate what remains uncertain.

Model calibration

Bayesian Workflow

The disciplined process: priors, prior predictive checks, sampling diagnostics, posterior predictive checks, and interval calibration (PIT).

Model calibration

Forecast Backtesting

Rolling-origin out-of-time accuracy, measured not asserted β€” including the negative control where the harness honestly fails.

For modelers

Real-Data Guide

What replaces the truth column on client data: prior→posterior learning, out-of-time backtests, stability checks, and interval honesty.

Experimental calibration

Calibration Loop

Wire experiments and the MMM into one self-correcting cycle β€” EIG, EVOI, three calibration mechanisms, and the virtuous loop.

Experimental calibration

Calibration Decisions

An interactive workshop: live calculators for the Bayesian update, EIG, EVOI in dollars, experiment portfolios, and geo-test power.

Experimental calibration

Continuous Learning

Learn channel response and synergies straight from designed geo experiments β€” no fitted MMM required, with a built-in stopping rule.

Foundations

Causal Inference

DAGs, do-calculus, and counterfactuals β€” the causal scaffolding behind every defensible MMM specification.

Methodology

Variable Selection

When to include precision controls, why confounders are non-negotiable, and how shopping for variables backfires.

For sponsors

For Business Stakeholders

The risks of specification shopping, the seven questions to ask any modeling partner, and what defensible measurement looks like.

See it run

Demos & Reports

Interactive workflow demonstrations, scenario analyses, and an example MMM report from a Q4 2025 fit.

Platform

Model Garden & Atelier

Author bespoke Bayesian models in-app, prove them against a compatibility suite, and publish versions any analyst can fit through the agent.

Reference

Technical Guide

Math and implementation details for every model in the framework β€” standard, nested, multivariate, and combined.

Reference · Sphinx

API Reference

Module-by-module API docs (Sphinx). Browse the source on GitHub or build locally with make html in docs/api.

Reference

Glossary

Definitions of every term used across the docs β€” adstock, EIG, EVOI, MFF, ITT, MDE, ROPE, and more.

Reference

FAQ

Common questions about MMMs, Bayesian methods, and the trade-offs the framework makes deliberately.

Project

Versioning & Changelog

Current version, API stability tier per module, versioning policy, and release notes.

Project

About

Background on the framework, its design principles, and the audiences it serves.

Ready for Measurement You Can Defend?

The framework is open source, the methodology is public, and every measured claim traces to a notebook you can re-run. Start with the documentation or go straight to the evidence.