Microsoft Clarity adds a Scrape-to-Referral ratio card ranking which AI operators take content and refer nothing back

Microsoft Clarity adds a Scrape-to-Referral ratio card ranking which AI operators take content and refer nothing back

analyticsAugust 14, 2026
By Antonio Fernandez

Microsoft added a Scrape-to-Referral ratio card to the Bot Analytics section of Clarity's AI Visibility Dashboard on 13 August 2026. The card reduces a long-running argument to one number per site: how much content a given AI operator crawls, measured against how many visitors that same operator actually refers back.

PPC Land reported the release on 13 August 2026. What follows sticks to what that report contains, plus clearly labelled analysis of what the change does and does not solve for site owners in Thailand.

What Microsoft shipped on 13 August 2026

Microsoft placed the new card inside Bot Analytics, which sits within Clarity's AI Visibility Dashboard. Alongside the ratio itself, the release carries operator-level ranked breakdowns that separate two groups: sources that send traffic, and sources that crawl heavily while referring almost nobody. That separation is the point of the feature. A raw crawl log tells you a bot visited. A referral report tells you a source sent people. Neither on its own tells you whether the exchange was fair.

The release bundles six components. There is the ratio card, the operator ranking, safeguards for mapped-domain coverage, handling for coverage gaps, click-through from the card into filtered session recordings, and traffic-quality analysis that goes past a raw referral count. The session-recording link is the part most likely to change day-to-day work, because it turns an aggregate number into a set of real sessions a marketer can watch.

What the ratio actually compares

The card expresses, as a single ratio, the volume of content an AI operator crawls from a site against the number of visitors that operator refers back to it. Read plainly, a high ratio says a large amount of a site's content was taken and very few humans arrived from the operator that took it. A low ratio says the exchange is closer to even.

Two things are worth stating carefully. First, the ratio is a measurement, not a verdict. An operator can crawl a lot and refer little because its product surfaces answers without links, or because a site's content is being used for training rather than retrieval, or simply because few users asked questions that page could answer. The card does not distinguish those cases and PPC Land's report does not claim it does. Second, the ratio is per site. That is what makes it useful. A number published by a third party describes somebody else's content mix, while this one describes yours.

The 6,000:1 figure is a product screenshot, not a benchmark

Microsoft's product illustration for the feature shows an average ratio of 6,000:1 against 41 total referrals. That figure comes from an illustration in Microsoft's own material. It is not an industry benchmark, it is not a median across sites, and nothing in the reporting presents it as a value any site should expect to match or beat.

This matters because screenshot numbers travel. A striking ratio lifted out of a product image and repeated in a slide deck becomes, within a few weeks, a figure people cite as if it were research. Anyone quoting 6,000:1 should say where it came from. The honest use of that number is as an example of how the card renders, not as evidence about the state of AI crawling.

The CDN requirement is the part that will block most sites

Populating bot data in Clarity requires a supported CDN connection. Microsoft lists Fastly, Amazon CloudFront and Cloudflare. Without one of those in front of the site, the bot data the card depends on does not arrive, and the card has nothing to display.

That single requirement decides who can use this feature and who cannot, and it has nothing to do with budget or skill. A site can have Clarity installed correctly, with sessions and heatmaps recording normally, and still see an empty Bot Analytics section because its traffic never passes through a supported CDN. The Clarity tag reports what happens in the browser. Bot traffic largely does not run a browser, so the crawl side of the ratio has to come from the edge, which is why the CDN link exists at all.

The six components, in one place

The table below lists what the 13 August 2026 release contains, using only the components named in the reporting. The first two rows are the visible surface. The rest is the machinery that decides whether the visible surface is trustworthy.

The six components, in one place
ComponentWhat it does
Scrape-to-Referral ratio cardExpresses content crawled by an AI operator against visitors that operator refers back, as one ratio
Operator-level ranked breakdownRanks operators and separates sources that send traffic from sources with heavy crawling and minimal referrals
Mapped-domain coverage safeguards and coverage-gap handlingGuards the numbers against incomplete domain mapping and gaps in coverage
Click-through into filtered session recordingsMoves from the aggregate figure into the individual sessions behind it
Traffic-quality analysisAssesses referred traffic beyond a raw referral count

How this fits the rest of Clarity's AI reporting

Microsoft published a separate blog post on 12 August 2026 that framed Clarity's AI Visibility set around grounding queries, citations and share of authority, plus competitive topic visibility. Read together with the ratio card, the two halves answer different questions. The 12 August framing is about whether a brand appears in AI answers at all. The 13 August card is about what an operator takes from the site while that is happening.

Most measurement work in AI search so far has sat on the visibility side, because that is the side that resembles rank tracking and is easier to sell. The cost side has been thinner. Putting a per-operator cost measure next to a per-operator visibility measure, inside the same free product, is a more complete picture than either alone.

What the reporting did not say

Several practical questions are open. The source did not state pricing or tier limits for the feature. It did not say whether the ratio is available through an API or an export, which decides whether the number can be pulled into a warehouse or a client report without manual screenshots. It did not describe how operators are identified, which matters because bot identification by user agent alone is easy to spoof and identification by verified IP range is not. And it gave no dates for any further rollout.

Those gaps are worth holding on to. A ratio you cannot export is a ratio that lives in one dashboard, and a ratio built on identification you cannot inspect is one you should sanity-check against server logs before making a blocking decision on it.

Why a free tool changes the argument

Until now the crawl-versus-referral imbalance was argued anecdotally, or through aggregate figures published by third parties about other people's sites. Both are weak footing for a decision. Anecdote invites the reply that your site is unusual. Aggregate third-party numbers invite the reply that your site is not in the sample.

A per-site, per-operator number inside a free analytics tool removes both replies. It also puts the measurement layer under a decision many site owners have already made without data. Blocking an AI crawler in robots.txt or at the edge is a trade: you give up whatever referral traffic and citation exposure that operator produces, in exchange for whatever content it stops taking. You cannot price that trade without knowing both sides of it for your own site. The ratio card is the first widely available tool that gives you the crawl side and the referral side in the same view.

What this means for Thai marketers

Clarity is free and is already installed on a large number of Thai sites, usually for heatmaps and session recordings. That makes this an unusual case: an AI-visibility measurement capability that a Thai SME can adopt without an enterprise analytics budget and without a new vendor contract.

The practical blocker is the CDN dependency, and it is worth spelling out before anyone opens the dashboard expecting numbers. A Thai site hosted locally, serving directly from its origin with no Fastly, CloudFront or Cloudflare in front, will see an empty card. Nothing is broken in that situation and no amount of retagging fixes it. The site simply has no edge layer reporting bot traffic into Clarity.

For a site that does sit behind a supported CDN, the sensible sequence is measurement before policy. Read the operator ranking for a full month before touching robots.txt, because a month is the shortest window in which a referral pattern separates itself from noise. Content-heavy Thai publishers and e-commerce catalogues will show the widest spread between operators, because they have the most content to crawl. Thin brochure sites will show very little either way, which is itself an answer.

One caution specific to bilingual sites. Thai and English versions of the same page are separate URLs, so crawl volume and referrals split across them. A ratio read at whole-site level can hide the fact that one language is being crawled and the other is being cited. If the dashboard allows it, look at the two language paths before drawing conclusions about an operator.

What to check this week

  • Confirm whether the site sits behind Fastly, Amazon CloudFront or Cloudflare. If not, the Bot Analytics section will stay empty and nothing else on this list applies yet.
  • Open Bot Analytics in the AI Visibility Dashboard and read the operator ranking before forming a view. Note which operators appear in the heavy-crawl, low-referral group.
  • Use the click-through into filtered session recordings on any operator that does refer traffic. Referral counts and referral quality are different things, and the recordings show which you have.
  • Compare the picture against your server or edge logs. The card is a reading, not a ruling, and a second source catches identification errors.
  • Write down the current numbers before changing any crawler rules. Without a baseline, you will not be able to tell what a block actually cost or saved.

None of this requires new spend. It requires reading the numbers for long enough to mean something, and it works alongside the analytics measurement setup a site already has rather than replacing it. Where a site is actively working on AI search visibility, the ratio belongs next to the rest of that programme, whether that is generative engine optimisation or the broader AI search work behind it.

Frequently asked questions

What is the Scrape-to-Referral ratio in Microsoft Clarity?

It is a single ratio comparing how much content an AI operator crawls from your site against how many visitors that operator refers back to it. Microsoft added the card to the Bot Analytics section of Clarity's AI Visibility Dashboard on 13 August 2026, together with an operator-level ranked breakdown that separates traffic-sending sources from heavy crawlers that refer almost nobody.

Do I need a CDN to see bot data in Clarity?

Yes. Populating bot data requires a supported CDN connection, and Microsoft names Fastly, Amazon CloudFront and Cloudflare. A site serving directly from its own hosting with none of those in front will see an empty Bot Analytics section even though the rest of Clarity keeps recording sessions normally.

Is the 6,000:1 figure an industry benchmark?

No. It appears in Microsoft's product illustration for the feature, shown against 41 total referrals, and it is a screenshot rather than a measured benchmark. Nothing in the reporting presents it as an average across sites or as a target, so quoting it as an industry figure would misrepresent it.

Can I export the ratio or pull it through an API?

The source did not say. PPC Land's report covers what the card shows and what it requires, but it does not state whether the ratio is available via API or export, and it does not give pricing or tier limits either.

Should I block AI crawlers based on this ratio?

Not on a first reading. The ratio gives you the cost side and the referral side of one operator for your own site, which is the measurement any block-or-allow decision should rest on, but a single week of data cannot separate a pattern from noise. Record a baseline, watch the operator ranking over a longer window, and check the figures against server logs before changing crawler rules.

If you want help reading these numbers for a specific site, or deciding whether an AI crawler is worth allowing, Relevant Audience works with Thai brands on exactly that question. Start with the measurement, then decide the policy.

Antonio Fernandez

Antonio Fernandez

Founder and CEO of Relevant Audience. With over 15 years of experience in digital marketing strategy, he leads teams across southeast Asia in delivering exceptional results for clients through performance-focused digital solutions.

Share to:
Copy link: