TL;DR
- PPC Land reported on 31 August 2026 on a preprint posted to arXiv on 21 August 2026 by Niklas Heusch, arXiv:2608.21128v1, on estimating marketing mix model parameters from geo-experiments.
- In the simulation a standard mix model reported 10.61x ROAS for paid search whose true return was 4.20x, with a 90 percent interval of 6.56 to 14.36 that missed the true value.
- Handed the real confounders that no practitioner possesses, the same model still returned 8.41x, so better controls alone do not close the gap.
- Structural estimation from four geo-experiments returned 4.14x with an interval of 3.79 to 4.48 that contained the true value.
- Everything is synthetic data from a single author with no peer review, so the 2.5 times figure is one simulation result and not a correction factor for any real account.
A preprint posted to arXiv on 21 August 2026 reports that a standard marketing mix model returned 10.61x return on ad spend for a paid search channel whose true return, inside the simulation, was 4.20x. That is an overstatement of roughly 2.5 times, and the model's 90 percent credible interval, running from 6.56 to 14.36, did not contain the true value at all.
PPC Land reported the paper on 31 August 2026. Before anything else about the finding: this is a single-author preprint running entirely on synthetic data, it has not been peer reviewed, and it does not measure any real advertiser's account. Every number below comes from a simulation.
What the paper is, and what it is not
The paper is titled "Structural Estimation of Marketing Mix Model Parameters from Geo-Experiments" and carries the identifier arXiv:2608.21128v1, filed under Applications in the statistics section, with a submission stamp of 21 August 2026. It lists one author, Niklas Heusch, and no institutional affiliation line. The contact address at the head of the paper resolves to a zalando.de domain, which PPC Land describes as the only signal in the document about where the work sits commercially.
A companion paper, cited as Heusch 2026, is credited with generating the synthetic dataset used throughout, and the paper states that the notebook producing that dataset is public.
So: no peer review, no audited advertiser data, no vendor tested by name, and a commercial context that has to be inferred from an email domain. Those are real limits and they belong at the top of the story, not in a footnote. What makes the paper worth reading anyway is that the claim it makes is structural rather than anecdotal, and it can be checked by anyone who runs the public notebook.
What a marketing mix model does, and why so many teams lean on one now
A marketing mix model takes aggregate weekly data, usually spend by channel plus sales, and estimates how much of the sales each channel caused. It does not follow individual users. It looks at totals over time and tries to attribute movement in the outcome to movement in the inputs.
That property is exactly why mix modelling came back. As user-level tracking degraded through cookie deprecation, consent requirements and app-level opt-outs, click-path attribution stopped covering the ground it used to cover. Aggregate modelling does not need identifiers, so it survived the change. Boards and finance teams now see mix model output in budget reviews, and channel budgets get moved on the strength of the numbers it produces.
Mix models translate spend into sales through two transformations. Adstock captures persistence, meaning how much of last week's spend still works this week. Saturation captures diminishing returns, meaning how quickly extra spend stops buying extra outcome. An effectiveness coefficient then converts the result into revenue. Every budget allocation an MMM recommends is a function of those three numbers.
What a confounder is, in plain language
A confounder is a third thing that moves both the input and the outcome, so the model credits the input with work the third thing did.
The everyday version in marketing is a promotion. A retailer runs a big December promotion, cuts prices, and raises paid search budgets at the same time because December is when people buy. Sales go up. Spend went up too. A model that does not properly account for the promotion and the season will hand some of December's demand to paid search, because paid search spend and December sales moved together.
Practitioners know this, which is why mix models include controls. The model in the paper had them. It was specified the way a competent practitioner would specify one: a promotion dummy, an observed promotional price level, and three annual Fourier harmonics for seasonality. Those are the standard seasonal controls in Robyn, Meridian and pymc-marketing. They were not enough.
The second result is the one that matters
The headline number will travel on its own: 10.61x reported against a true 4.20x. The comparison is only possible because everything happens in a simulation where every parameter is known. There is no observational dataset anywhere in which the true return is written down.
The result underneath it is more useful. The same model was then handed the actual confounders from the data-generating process: the true promotion state, the true price level, the latent seasonal component, product-quality drift, and market sentiment. No practitioner possesses those covariates. Nobody has them, in any account, ever. With all of them, the model still returned 8.41x, with an interval of 6.91 to 9.81 that also excluded 4.20.
The paper is explicit that this specification exists as a diagnostic and not as a proposal. Its job is to bound what better controls could ever achieve, and the bound is roughly double the true return. PPC Land's framing of the point is the one worth keeping: this is a different claim from the familiar one about data quality, because it says the gap is not a measurement problem that a richer dataset closes.
That distinction is the whole story. "Our data is not good enough" implies a fix: collect more, clean it, add covariates, buy a better feed. "This class of model cannot separate these effects" implies something else, that adding data to the same specification moves the estimate from very wrong to somewhat less wrong and then stops.
What the paper proposes instead
The paper argues that geo-experiment time series can recover the correct figure, and that standard practice throws away most of what an experiment produced.
A geo-experiment divides a market into geographic units, typically Designated Market Areas or metropolitan regions, and assigns them at random to treatment or control. During the test window, spend in the treatment group is modified, often reduced to zero in what the industry calls a go-dark test, while the control group carries on spending normally. A well-designed test has three phases: a pre-test period establishing that both groups follow parallel trends, a test period in which spend differs, and a cooldown period in which spend returns to normal and residual effects decay.
Both groups produce complete daily time series of outcomes and spend across all three phases. Standard practice then collapses all of it into a single number: total incremental revenue over the test period divided by total incremental spend. The paper's objection is that summing and dividing discards the information about dynamics, meaning how effects build and decay across weeks and how they respond to different spending levels.
The results table
PPC Land published the paper's comparison of four methods against a known truth of 4.20x.
| Method | ROAS | 90% credible interval |
|---|---|---|
| True value in the simulation | 4.20x | Known, not estimated |
| MMM, realistic controls | 10.61x | 6.56 to 14.36, excludes the truth |
| MMM, oracle controls | 8.41x | 6.91 to 9.81, excludes the truth |
| Structural estimation, 2 tests | 4.31x | 3.87 to 4.75, contains the truth |
| Structural estimation, 4 tests | 4.14x | 3.79 to 4.48, contains the truth |
One detail from the underlying parameters is worth pulling out, because it has an operational consequence a ROAS figure hides. The true adstock decay was 0.20, meaning 20 percent weekly carryover. The observational model estimated 0.50. A model that believes half of last week's spend carries into this week describes a channel that can be pulsed, flighted and rested. A channel with 20 percent weekly carryover cannot. Those are different media plans, produced from the same data.
The limits, stated plainly
Everything in the empirical section is synthetic, and the paper says so. There is no validation against a real advertiser's sales, no holdout on live campaign data, and no comparison with an audited outcome. The reason is structural rather than lazy: ground truth for advertising effectiveness does not exist in observational data, so the only place an estimator can be checked against a known true return is a simulation where somebody wrote the true return in the first place.
The cost of that is that the demonstration inherits every assumption in the simulation, including the assumption that the real process takes the exact functional form the estimator assumes. The factor of 2.5 is a property of this simulation. It is not a constant, and it does not mean every mix model overstates every channel by 2.5 times.
The source also does not include a response from Google or from any mix modelling vendor. The paper tests a specification the paper describes as what a practitioner would actually run, using the standard seasonal controls of three open-source frameworks. It does not test any commercial product, and PPC Land is careful about that distinction.
Nothing in the paper or the report concerns Thailand, Thai advertisers or any Southeast Asian market.
The market this lands in
PPC Land sets the paper against two years of rebuilding in aggregate measurement. Google opened Meridian globally in January 2025 after testing with hundreds of brands, added non-media variables, channel-level contribution priors and binomial adstock decay in September 2025, launched a Scenario Planner in February 2026, and announced Meridian GeoX and Meridian Studio on 5 May 2026. GeoX is a geographic incrementality testing tool whose stated function is to feed experiment results into the mix model.
The Heusch paper is, in effect, an argument about how that handover should work. Feeding a single lift number into a model is not the same as feeding the experiment's full time series into it, and the paper's claim is that the difference is large.
What a marketer should actually do with this
Treat mix model output as a hypothesis, not a measurement. A number produced by a model that cannot see the confounders is a starting point for a test, and it deserves the same scepticism as any other estimate that arrived without an experiment behind it.
Run incrementality tests on the channels carrying the most budget, and compare the experimental result against what the model said. Where they disagree by a wide margin, the model is the thing under suspicion. Holdouts do the same job at lower cost when a full geo test is out of reach.
Ask whoever owns the model two questions: which confounders does it control for, and which ones does it not. The second answer is the one that predicts where the estimate will be wrong. If the answer to the first question is only seasonality and promotions, that is precisely the specification the paper tested.
Watch for the channels most exposed to feedback loops. Paid search sits inside automated bidding that raises spend when conversion rates rise, and conversion rates often rise because demand is strong rather than because the ads worked. That is the mechanism the simulation was built around, and it is the default operating condition of the channel. Anyone reporting on Google Ads performance should know how their reported return was produced before defending it in a budget meeting.
Keep the measurement plumbing honest as well. A model fed by inconsistent conversion definitions or a broken data layer will be wrong for reasons that have nothing to do with confounding, and those failures are fixable. That is ordinary work under analytics and GA4 implementation, and it should be done before anyone argues about model specification.
What this means for Thai marketers
The paper says nothing about Thailand and no Thai advertiser appears in it. The relevance is indirect and it is about reporting practice rather than about a local finding.
Media budgets in Thailand are reported to clients and boards with ROAS attached, and a growing number of those figures come from aggregate models rather than from tracked conversions. If a mix model can overstate a paid search channel by 2.5 times under textbook controls, then a reported ROAS deserves a note about how it was produced. That is a defensible position to take with a client, and it is more honest than presenting a modelled number as a measured one.
The cost side is worth being realistic about too. A four-week go-dark test on a performance channel across a meaningful share of a market is a real revenue line, and the paper's method needs several such tests at different spending levels. Most advertisers run one or two incrementality studies a year, which is below what the method needs. For smaller accounts, a holdout region and a careful before-and-after read is what is actually affordable. Organic work under SEO in Thailand has the same problem in a different shape, since organic contribution is usually inferred rather than tested at all.
Frequently asked questions (FAQ)
Does this mean every marketing mix model overstates ROAS by 2.5 times?
No. The 2.5 times figure comes from one simulation under one model specification, and the paper does not claim it generalises. It is evidence that a standard specification can be badly wrong in a known direction, not a correction factor to apply to anyone's reported numbers.
Has this paper been peer reviewed?
No. It is a preprint posted to arXiv on 21 August 2026 under the identifier arXiv:2608.21128v1, with a single listed author and no institutional affiliation line.
What were the oracle controls, and why do they matter?
The oracle controls were the true confounders taken directly from the simulation's data-generating process, including the true promotion state, the true price level, the latent seasonal component, product-quality drift and market sentiment. They matter because no practitioner has them, so the 8.41x the model produced with them bounds what better data could ever achieve.
Is a geo-experiment the answer for every advertiser?
Not at any budget. The method needs tests long enough to show carryover dynamics and spread across different spending levels, which is more than a single lift study, and going dark on a performance channel costs real revenue while the test runs.
Does the paper say anything about Thailand or Southeast Asia?
No. The paper and PPC Land's report concern a simulated online retailer with three channels and make no reference to Thailand, Thai advertisers or any regional market.
If a reported ROAS in your account has never been checked against an experiment, that is a reasonable thing to fix before the next budget cycle, and it starts with knowing how the number was produced.







