How to measure GEO, and what happens if you stop

How to measure GEO, and what happens if you stop

geoAugust 18, 2026
By Antonio Fernandez

Measuring GEO takes two separate layers of measurement, because no single tool reports it. The first layer is how often an AI assistant mentions your brand in its answer, which you collect by running the same set of prompts on a schedule and logging the results by hand. The second layer is how many people read that answer and clicked through to your site, which GA4 shows you but always undercounts. Nothing gives GEO an impression count the way Google Search Console gives one to SEO. This article covers what each method tells you, what it misses, how to set a baseline before you start, and what happens to the work if you stop.

What GEO is and how it differs from SEO

GEO stands for Generative Engine Optimization, the work of getting your content used inside answers produced by language tools such as ChatGPT, Perplexity, Gemini, Microsoft Copilot, and the AI Overviews box Google shows above ordinary results. The outcome you want is not a rank position. It is presence inside the answer itself, either as a sentence that names your brand or as a citation link shown beside the answer. For the full service view of this work, see the GEO service page.

GEO vs SEO differ in the unit of the result

SEO has units the whole industry agreed on: position, impressions, and clicks, all reported by one shared source that Google itself publishes. Site owners and agencies can argue over the same numbers. GEO has nothing equivalent. ChatGPT and Perplexity count no impressions for anyone, publish no dashboard telling you how many answers named your brand this month, and produce answers that shift with the model version, the language of the question, and the history attached to the account asking.

The shape of a win differs too. In SEO one page beats competing pages for one query. In GEO you win on the consistency of the facts about your brand across every place they appear. If your site lists one price band, a directory lists another, and your company profile lists a third, a system that has to summarise an answer will lean on whichever source set agrees with itself, not the one that reads best.

Which businesses GEO suits, and which it does not

GEO pays off most for businesses whose customers research heavily before deciding, because people ask AI assistants questions during comparison, not at the moment of payment. Someone typing a question into ChatGPT about ERP options for a mid-size factory in Thailand sits at the front of a buying process that runs for months. Businesses sitting at that point benefit right away.

The categories that clearly qualify:

  • Software and subscription services, where buyers compare features and pricing before speaking to sales
  • Professional services such as accounting firms, legal advisers, and tax consultants, where the first question a buyer asks is a knowledge question
  • Industrial manufacturers and distributors that already publish numeric specifications and technical documents
  • Schools, hospitals, and specialist clinics, where people ask about eligibility, process, and cost in advance
  • Any business holding data nobody else has, such as a published rate card, its own product test results, or internal statistics it is willing to release

The businesses where GEO usually fails to repay the effort share one trait: a short purchase decision that never passes through research. A neighbourhood food shop, a convenience store, a local barber. Those customers walk past and come in, or search a map when they are hungry. Rewriting content so an AI can quote it barely moves the revenue. A map profile and a review programme are worth more.

Another group worth pausing on: brands whose sales come from seeing something and wanting it immediately, such as streetwear or a cafe growing on short video. The channel creating demand there is the video feed, not the answer box. New companies nobody has mentioned anywhere cannot force GEO either, because a summarising system wants more than one source confirming a fact, and your own website alone is not enough for a model to risk naming you. That case calls for getting mentioned first and doing GEO after.

Then there is time and budget. A business that needs leads inside three weeks to hold its cash position should not put the first money into GEO, because the refresh cycle of the AI tools is not under anyone's control on your side. That budget belongs in paid media first, with the patient portion of the budget going to the longer work.

How to do GEO SEO so your content gets used

GEO work splits into three layers: the technical layer that lets bots reach the content, the content structure layer that makes passages easy to lift, and the off-site layer that keeps the facts about you consistent. All three overlap heavily with SEO, but the details at the far end differ.

Technical layer: let the AI bots actually reach the page

Start with robots.txt and check you are not blocking the relevant crawlers: GPTBot and OAI-SearchBot from OpenAI, PerplexityBot from Perplexity, and ClaudeBot from Anthropic. Google-Extended is a separate control governing use in the Gemini app and Google grounding products, and it does not govern appearance in AI Overviews, which follows ordinary Googlebot access. Plenty of sites blocked these crawlers in 2023 and forgot, so the block has to come off before anything else matters.

Rendering comes next. If your main content is drawn entirely by client-side JavaScript, a bot that fetches raw HTML gets an empty page. The quick check is to open the page source and search for a sentence from your first paragraph. If it is not in the HTML, the fix is server-side rendering or static generation. That check sits alongside page speed and index status in a single pass during an SEO audit.

Content layer: write passages that survive being lifted out

The rule that holds up is that every subheading answers its question in the first sentence, then expands. Systems that pull text into an answer tend to take a paragraph or a short block. If your paragraph opens with two lines of throat-clearing, the block that gets lifted carries nothing.

Next, each sentence has to stand on its own. Avoid backward references like "as mentioned above" or "this method", because once the sentence is lifted off the page the reader has no way to know what it points at. Name the actor inside the sentence. Write that Google counts AI Overviews data inside the Web totals of the performance report, rather than writing that the data is counted in the main totals.

The formats that get reused most often are comparison tables with real values in the cells, ordered step lists, and question blocks phrased in the words people actually type rather than the words a marketer assumes they type. For how this maps to conversational question behaviour specifically, the ChatGPT SEO page goes deeper on that side.

On schema, ship at least three types: Organization with sameAs pointing at your official profiles, Article for editorial pages, and FAQPage for question blocks. Schema does not oblige any system to pick you. It makes the facts about your identity readable without guesswork.

Off-site layer: make the facts agree everywhere

Inventory every place your brand data lives: business profiles, industry directories, profile pages on trade media, and your own about page. Then make the company name, founding year, address, service scope, and price band match at every one. Conflicting data is the reason a system chooses to say nothing about a brand rather than guess. This work belongs to AI SEO directly, because it manages your brand identity as machines read it rather than adding more articles.

How to measure GEO

Honestly measuring GEO performance starts with accepting that no single number covers it. What works is assembling the picture from four sources that each see a different angle: a manual prompt tracking sheet, GA4, Search Console, and your server logs, topped up by asking buyers on your forms where they heard about you.

The table below sets out what each measurement tool sees and what it cannot see. Before deciding what to report on, it helps to know the boundary of each one.

How to measure GEO
MethodWhat it tells youWhat it misses
Manual prompt tracking sheetHow many runs mentioned your brand, which competitors appeared in the same answer, and which domains were cited as sourcesHow many real people ask that question
GA4 session sourceSessions whose referrer is an AI assistant domain, post-click behaviour, and conversions from those sessionsClicks that arrive with no referrer and land in Direct
Google Search ConsoleQueries gaining impressions while CTR falls, and the trend in branded query impressionsAI Overviews figures broken out separately, and anything from ChatGPT or Perplexity
Server logsWhich pages GPTBot, OAI-SearchBot, and PerplexityBot fetched, and how oftenWhether a fetched page was actually used in an answer
How did you hear about us fieldBuyers who remember meeting the brand through an AI assistant and say soBuyers who do not remember, guess, or skip the field

Set a baseline before you start

Capture the baseline before touching any content, because without it you will never separate the month-three result from a model version update that happened to land at the same time. Four things get recorded on day one: the results of the full prompt set you plan to track, the count of sessions from AI assistant domains in GA4 over the previous 90 days, the impression count for branded queries in Search Console over the previous 90 days, and the list of pages AI crawlers have already fetched according to your logs.

The prompt set should run roughly 20 to 40 questions in three shapes: category questions with no brand name in them, comparison questions that name a competitor in the prompt, and direct questions about your brand to see whether the system describes you correctly or wrongly. A business selling to both Thai and international buyers needs two sets split by language, because Thai-language answers and English-language answers draw on different source pools.

The manual prompt tracking sheet

A manual prompt tracking sheet is the most direct GEO measurement available right now. Open a spreadsheet and set these columns: test date, prompt text, tool used, question language, brand mentioned yes or no, where in the answer the mention fell relative to other names, competitor names appearing in the same answer, the domains cited as sources, and whether a link back to your site was present.

The most common mistake is running each prompt once and drawing a conclusion. Model answers are not stable. Ask the same question three times and you may get three different lists. What works is running each prompt three times and recording a fraction, such as mentioned in 2 of 3, then rolling the whole set into one figure: this month the brand was mentioned X times across Y runs. That figure is the headline metric you compare month over month.

Watch the account you test from. If you are logged into an account that has discussed your own brand dozens of times, a system that carries user context may surface your name because of that history rather than because of your content. Test in temporary chat mode or on an account with no personal history, and note which country you tested from, because answers change with location.

Fortnightly or monthly is the right cadence. More often than that and you are watching model noise. Less often and you miss the moment a provider switches its default model. Record the model version every time, because when the numbers fall off a cliff, the usual explanation is a changed default model rather than content that suddenly got worse.

What GA4 shows and where it misses

GA4 sees AI assistant traffic as referrals. Open the traffic acquisition report, switch the dimension to session source, or build an exploration filtered to sources containing chatgpt.com, perplexity.ai, gemini.google.com, and copilot.microsoft.com, then save it as a reusable comparison. The advantage of this route is that you also see post-click behaviour: landing page, time on site, and conversions, which tells you whether visitors arriving from AI answers behave differently from ordinary search visitors.

Where GA4 misses is that many clicks from AI assistants arrive with no referrer attached, whether from in-app browsers, from a user copying the link and pasting it, or from someone reading the answer and searching your brand name later. All of that lands in Direct or back in Organic Search. The AI referral number in GA4 is a floor on the truth, never the truth.

What Search Console can and cannot tell you

Google states that AI Overviews and AI Mode data are counted inside the Web search type totals of the performance report, with no filter to break them out. You see the combined total but cannot isolate the share coming from the answer box. What you can do is read the pattern: if a group of queries gains impressions while CTR falls steadily, that is consistent with an answer being shown above the results. It is an inference, not a measurement.

What Search Console cannot tell you is anything that is not Google Search. Traffic and citations in ChatGPT, Perplexity, Claude, or Copilot never appear in that report in any form. Anyone reporting GEO results with Search Console screenshots alone is reporting on a different channel.

Search Console does have one angle that helps GEO specifically: impressions on branded queries. When people meet a brand inside an AI answer, a portion of them do not click the citation and search the brand name instead. If branded query impressions climb during a GEO programme with no advertising or press push running, that is a usable indirect signal, even though it proves no causal link on its own.

Why GEO attribution is genuinely hard

The core problem is that the place the effect happens and the place you measure sit apart. The answer is produced inside the provider's app, which sends nothing back to the site owner. Plenty of users get a complete answer and close the tab, leaving no click to count. The interested ones often search the brand name as a second step, and by the time they fill in a form the path has crossed several days and several channels. A last-click attribution model hands the credit to whatever came last.

Two workarounds hold up. The first is a "how did you hear about us" field on the contact form with ChatGPT or AI assistant among the options. That data is incomplete, but it is the one layer that comes from the buyer directly. The second is accepting trend measurement instead of per-lead measurement: compare the three months before you started with the three months after, and read three lines together, being the mention rate in the prompt sheet, AI referral traffic, and branded query impressions.

The next question everyone asks is what counts as good. This has to be said plainly: no reliable public benchmark exists for AI citation rates. No organisation publishes industry medians, and the figures circulating online come from prompt sets each measurer chose themselves, which makes them impossible to compare across. The only usable benchmark is your own first month. If the brand was mentioned 8 times across 90 runs in month one, that is the baseline, and next quarter's goal is to move that number rather than chase someone else's.

An infographic summarising GEO measurement from four sources: a manual prompt tracking sheet counting how many runs mentioned the brand, GA4 sessions from AI assistant domains, Search Console queries gaining impressions while CTR falls, and server logs showing which pages AI bots fetched, with a footer noting that each prompt must be run three times and that no industry benchmark exists so your own first month is the baseline.

What happens if you stop doing GEO

Stopping GEO work does not erase the results overnight. The results get gradually crowded out by newer content instead. What happens splits into what decays and what stays.

The first thing to decay is prompt coverage. Questions where your brand used to be named get taken over question by question by competitors who kept publishing. The second is freshness. Pages carrying stale prices, product versions, or service terms get skipped more often once a system finds a source whose data matches the present. The third is off-site corroboration that stops growing. The profiles and mentions do not vanish, but with nothing new added the combined weight falls relative to competitors still moving.

What stays is nearly all of the structural work. Moving to server-side rendering, opening robots.txt to the crawlers, adding schema, and aligning your name, address, founding year, and service scope across every source are one-time jobs that keep holding. Content you already published stays indexed and can keep being cited, particularly pages carrying data available nowhere else, such as a published rate card or your own test results. Those pages decay slowest because there is no substitute source for the system to pick instead.

The timescale needs saying clearly: no reliable public data quantifies the decay rate of GEO results. Anyone telling you that three months off costs half your visibility is describing something the industry has not measured. What can be said from the mechanics is that tools searching the live web at answer time, such as Perplexity and AI Overviews, reflect changes on the web faster, while whatever is baked into model weights only shifts when a new version is trained, on a schedule no site owner controls.

The practical suggestion, if you are stopping: stop the content production but keep the prompt sheet running. Run the same set for another three to six months while changing nothing else, and watch how much the mention rate drops per month. That is your own decay curve, and it is worth more to next year's budget decision than an industry average that does not exist. If you restart later, the restart costs less than the first round did, because the technical layer is already in place.

How to choose a GEO agency

Ask four questions before you look at the quote. First, what will you measure this with? If the answer is a keyword ranking report and nothing else, that is SEO work under a new name. Second, how many prompts are in the tracked set, who chooses them, and how many runs per prompt? Third, will you capture a baseline before starting, and what goes into it? Fourth, if month three shows no movement, what is the plan? A good answer talks about changing the content set and adding off-site corroboration rather than only asking for more time.

One more question worth asking: who touches the website code? Half of GEO is technical work that has to be fixed in the system. If the agency can only hand over a recommendations document for your team to implement, the real cost runs higher than the quote shows. Teams who have done SEO in Thailand before tend to have the advantage of telling content problems apart from rendering problems and talking to developers without a translator in the middle.

If you want help judging whether your business fits the conditions where GEO pays off, and what baseline to capture before starting, get in touch and we will assess whether it is worth the effort yet.

GEO measurement FAQ

Can GEO be measured as ROI?

It can be estimated, but not measured directly the way paid advertising is. The closest approach is conversions from sessions whose referrer is an AI assistant domain in GA4, plus the leads who select ChatGPT or AI assistant in a "how did you hear about us" field. That total sits below reality, because referrer-less clicks and people who searched your brand name afterwards are counted in neither path.

How many prompt runs before the numbers mean anything?

Run each prompt at least three times, because model answers are unstable and the list of names can change on every run. A 30-question set therefore means 90 runs per measurement cycle, reported as a fraction such as mentioned in X of 90. Running a prompt once and drawing a conclusion is the most common mistake in GEO measurement.

Can I see ChatGPT traffic in Search Console?

No. Search Console reports only Google Search data, so traffic and citations from ChatGPT, Perplexity, or Copilot appear nowhere in it. Use the GA4 traffic acquisition report filtered on session source instead. Google's own AI Overviews data is folded into the Web totals with no separate filter.

What citation rate counts as good?

No reliable benchmark exists for that question. No public source publishes industry medians for AI citation rates, and the figures circulating come from prompt sets each measurer picked, which cannot be compared across businesses. Use your own first month of measurement as the baseline and judge progress against that line.

If I stop GEO for three months, does it all disappear?

Not all of it. The parts depending on freshness get replaced first, while pages you already published stay indexed and can still be cited, and technical work such as schema and server-side rendering stays in place. What goes fastest is the questions competitors who kept publishing take over. No public data quantifies the decay rate, so the only way to know yours is to keep running the same prompt set while you are stopped.

Antonio Fernandez

Antonio Fernandez

Founder and CEO of Relevant Audience. With over 15 years of experience in digital marketing strategy, he leads teams across southeast Asia in delivering exceptional results for clients through performance-focused digital solutions.

Share to:
Copy link: