Google's Meridian GeoX leaves beta with an unreplicated 31% savings claim

Google's Meridian GeoX leaves beta with an unreplicated 31% savings claim

analyticsSeptember 11, 2026
By Antonio Fernandez

TL;DR

  • Google's Meridian GeoX, an open-source geographic incrementality testing library, reached general availability on 9 September 2026, announced on a Google Analytics YouTube livestream.
  • Its product manager claimed budget savings of more than 31% for large advertisers versus unnamed open-source alternatives, and design generation times cut by over 94%.
  • PPC Land reported that no sample size, test period, comparison set or definition of a large advertiser was disclosed, and that neither figure has been independently replicated.
  • Meridian 2.0.0 shipped alongside it, making JAX the default backend in place of TensorFlow, with claimed gains of twice the speed and four times the memory efficiency.
  • Nothing was stated about Thailand, Thai availability, pricing or a minimum spend level.

Google moved Meridian GeoX out of beta and into general availability on 9 September 2026, announcing the release during a livestream on the Google Analytics YouTube channel. GeoX is an open-source library for geographic incrementality testing, and the Google product manager who presented it claimed budget savings of more than 31% for large advertisers compared with unnamed open-source alternatives.

What Google released on 9 September 2026

PPC Land reported on 9 September 2026 that Meridian GeoX had exited beta and reached general availability, and that the announcement was made during a livestream on the Google Analytics YouTube channel. The full report is at PPC Land. Two Google product managers presented on the stream. Katie Monroe, product manager for the core Meridian library, covered the wider release. The GeoX product manager presented the testing library itself, and the automatic captions on the stream rendered that person's name inconsistently, so this article does not attach a name to those remarks.

According to that 9 September 2026 report, GeoX measures advertising uplift by comparing treatment geographic groups against control geographic groups. The library uses stratified sampling and time-based regression, and it handles multicell designs natively, which means more than one treatment group can run against a control inside a single experiment rather than as a series of separate tests.

PPC Land reported that the same release carried Meridian version 2.0.0, dated 2 September 2026 in the public changelog. In 2.0.0, JAX replaced TensorFlow as the default computational backend, with claimed improvements of twice the speed and four times the memory efficiency. The same report said time-varying media ROI is planned for the fourth quarter of 2026.

What geo incrementality testing measures

This section is analysis, not reporting. Most advertising measurement answers a bookkeeping question: of the conversions that were recorded, which touchpoint gets the credit? Last-click attribution, data-driven attribution and modelled conversions are all different rules for splitting credit across events a measurement system already observed. None of them tell you what would have happened if the advertising had not run. That second question is the one a finance director tends to be asking, and it is a causal question rather than an accounting one.

Geo incrementality testing exists because that causal question got harder to answer at the user level. Cookie restrictions in browsers, app tracking permissions, consent requirements and the growth of modelled data all reduce how much of a single person's path a platform can observe end to end. Geography survived all of that. A region is not a cookie, it does not opt out, and it does not disappear when a browser clears storage. So a holdout region became one of the few clean reads left: the advertiser deliberately changes spend in some places, leaves it alone in others, and compares.

How a treatment and control split works, step by step

  1. Split the market into geographic units that can be targeted separately, such as provinces, metro areas or postcode groups.
  2. Group those units by how they behave at baseline, so that high-volume and low-volume units are represented on both sides rather than clustered on one side.
  3. Assign some units to treatment and the rest to control.
  4. Change the advertising in the treatment units only. That change can be turning spend on, turning it off, or moving it up or down by a set amount.
  5. Hold everything else steady for the length of the test. New creative, a price change or a promotion that lands in the middle of the window will show up in the result and will look like advertising effect.
  6. At the end of the window, compare the outcome observed in the treatment units against the outcome the control units imply would have happened without the change.

The control group is doing one job in that sequence: standing in for a version of the treatment regions where the advertising change never happened. Everything that makes a geo test credible or worthless comes down to how well it does that job.

What a multicell design adds

A single treatment group against a single control gives one number: the uplift at one spend level. A multicell design runs several treatment groups at different spend levels at the same time, against a shared control. That returns a shape rather than a point, which is what a budget decision actually needs. Knowing that a channel was incremental is less useful than knowing where the incremental return starts to flatten. PPC Land reported that GeoX supports multicell designs natively, which places that capability inside the library rather than in the analyst's spreadsheet.

Matched markets and why out-of-sample validation is the harder test

PPC Land described stratified sampling with out-of-sample validation across historical periods as the methodology that separates GeoX from matched markets approaches, which the outlet said rely on in-sample fit. The distinction is worth unpacking, because it decides how much weight a geo test result can carry.

A matched markets approach picks control regions that resemble the treatment regions historically, then judges the pairing by how closely the model reproduces the history it was built on. That is in-sample fit. The problem with in-sample fit is structural: a model with enough flexibility can be made to reproduce almost any past series, and a pairing tuned until the historical lines overlap has been fitted to noise as well as to signal. A good in-sample fit is therefore weak evidence that the pairing will predict anything.

Out-of-sample validation withholds periods from the fitting process, builds the pairing on what is left, then checks whether it predicts the withheld periods it never saw. A pairing that survives that has demonstrated something the in-sample version has not. It is the same discipline a forecaster or a machine learning practitioner applies for the same reason, and it is a harder bar to clear.

None of this makes a geo test automatically correct. Out-of-sample validation tests whether the control tracked the treatment before the experiment. It cannot tell you that something unrelated to advertising hit one group during the test window.

The claimed figures and what was not disclosed

The release came with three headline numbers. The table below sets each claimed figure against what it was measured relative to, and what PPC Land reported as absent from the disclosure.

The claimed figures and what was not disclosed
Claimed figureMeasured againstNot disclosed
Budget savings of more than 31% for large advertisersUnnamed open-source alternativesSample size, test period, comparison set, and the definition of a large advertiser
Design generation times cut by over 94%Other open-source and legacy solutionsWhich solutions, on what hardware, at what design complexity
Twice the speed and four times the memory efficiency in Meridian 2.0.0JAX as the default backend in place of TensorFlowWorkload, model size, and benchmark method

PPC Land put the caveat plainly: the figures rest on "independent replication that has not yet happened". That framing is correct and it is the part most likely to be dropped when a number like 31% travels through a marketing deck. A vendor benchmark against unnamed competitors, with no sample size and no stated test period, is a claim about a product. It becomes a result when someone outside the vendor runs it and gets the same answer.

There is a second reason to hold the 31% figure loosely, and this is analysis. A budget saving from a measurement tool is not a measured quantity in the way a conversion count is. It depends on what the advertiser would have spent under the alternative, which means it depends on a counterfactual budget decision by a human. Without the definition of a large advertiser and the comparison set, there is no way to tell whether the saving describes better experiment design, a different default setting, or a different assumption about what the baseline would have been.

What an advertiser needs before a geo test is possible at all

This is analysis. A geographic experiment has requirements that a click-level report does not, and an account that misses them will produce a result with confidence intervals wide enough to cover almost any conclusion.

  • Enough geographic units. The unit of observation in a geo test is the region, not the click. Splitting a country into four regions gives four observations. Statistical power comes from the number of units, so a market that can only be cut into a handful of targetable areas limits what any library can recover.
  • Enough spend and volume per unit. A region that produces a small number of conversions per week carries noise large enough to swallow a real uplift. The effect has to be big relative to the week-to-week variation in that region.
  • A stable baseline. A business in the middle of a launch, a repricing or a distribution change has a moving baseline, and the test will attribute that movement to the advertising.
  • Operational discipline. Someone has to be able to say no to a mid-test promotion, a creative refresh or a bid strategy change in the treatment regions for the full window.
  • Tolerance for the cost of the holdout. If the test works by switching advertising off somewhere, the business gives up whatever that advertising was earning in those regions for the duration. That is the price of the answer.

Those constraints are why geo testing has been the preserve of larger advertisers, and why the claimed savings were framed around large advertisers rather than around every account.

What this means for marketers in Thailand

PPC Land's report said nothing about Thailand. There was no statement about Thai availability, no pricing, and no minimum spend threshold. The report described Meridian as open source, and it named no regional restriction, but the absence of a restriction in an article is not the same as a confirmation, so treat Thai availability as not stated.

There is a structural point worth raising for a domestic campaign, and it is analysis rather than anything reported. Commercial activity in Thailand is concentrated heavily in and around Bangkok, with a long tail of provinces that individually generate small volumes. That concentration makes geographic cells hard to balance: one cell carries most of the demand, and the rest carry noise. An advertiser considering a geo test on Thai traffic should look first at how many targetable regions actually produce a workable weekly conversion count, because that number, not the library, sets the ceiling on what the test can detect.

A more immediate use for most accounts is upstream of any experiment. Geo testing rewards clean, stable measurement, and an account whose conversion definitions drift or whose regional reporting is inconsistent cannot support the test in the first place. Getting analytics and conversion tracking into a trustworthy state is the prerequisite, and Relevant Audience covers that work in its GA4 migration and analytics setup service. The regional performance data that a geo design needs also has to come out of the ad platform in a usable shape, which sits inside day-to-day Google Ads management.

Frequently asked questions

Is Meridian GeoX available in Thailand?

The source did not say. PPC Land's 9 September 2026 report covered general availability without naming any country or region, and it did not state whether availability differs by market. Meridian is described in the report as open source, which usually implies no geographic gate on the code itself, but no statement about Thailand appeared in the article.

Does GeoX cost anything to use?

PPC Land did not report a price or a minimum spend requirement. The report described Meridian as an open-source library. Any cost of running a geo experiment sits in the advertising that gets switched off or moved during the test, and in the analyst time to design and read it, rather than in a licence fee that the report mentioned.

Do I need to do anything if my team already uses Meridian?

Nothing in the report placed a requirement on existing users. As analysis, the change most likely to matter operationally is in version 2.0.0, where JAX replaced TensorFlow as the default backend, since a backend swap in a modelling library is the kind of change worth reproducing a known result against before trusting new output.

Should I plan a budget around the 31% saving?

No, because it is a vendor figure that has not been independently replicated. PPC Land reported that the claim came with no sample size, no test period, no named comparison set and no definition of a large advertiser. A number in that condition is a marketing claim about a product rather than a measured outcome you can forecast against.

Does a geo test replace the conversion data in my ad account?

No, they answer different questions and both stay useful. Platform conversion data tells you which recorded events followed which interactions, at a level of detail a geo test cannot reach. A geo experiment estimates how much of the outcome would have happened anyway. Neither one is a substitute for the other, and the second is usually the one a budget argument turns on.

Where this leaves measurement teams

The honest summary of 9 September 2026 is that a well-regarded open-source measurement stack added a geo experimentation module and put two unreplicated efficiency numbers on it. The method PPC Land described, stratified sampling with out-of-sample validation, is the stronger of the two approaches it was contrasted with, and that part does not depend on anyone accepting the 31% figure. Advertisers with enough regions and enough volume to run geo tests have another free tool for it. Everyone else is better served by fixing the measurement that a geo test would sit on top of. If you want a read on whether your account and your analytics could support incrementality testing at all, Relevant Audience is happy to take a look.

Antonio Fernandez

Antonio Fernandez

Founder and CEO of Relevant Audience. With over 15 years of experience in digital marketing strategy, he leads teams across southeast Asia in delivering exceptional results for clients through performance-focused digital solutions.

Share to:
Copy link:

Read us often? Add Relevant Audience as a preferred source so our articles surface more in your Google results.