TL;DR
- A census of 4,776 operating restaurants, cafes and bars in greater Canggu and greater Ubud found 4,087 of them, 85.6 percent, were never recommended across 2,208 AI runs.
- Owning a website carried the largest entry effect at an odds ratio of 1.92, while average star rating was null at entry, 0.89, and only predicted first position, at 1.17.
- The four systems recommended permanently closed venues 93 times across 14 establishments, while outright fabrication was 0.08 percent of valid mentions.
- Repeating an identical query on the same engine gave mean top-20 overlap of just 0.22 to 0.45, so a single-run visibility check measures noise.
- Vladimir Pitenin of Norly Research posted the paper, arXiv:2608.07069, on 7 August 2026; Norly funded and conducted it and discloses the competing interest.
Four AI assistants were asked 2,208 times where to eat in two parts of Bali, and 85.6 percent of the local food and drink market never came up once. A paper posted to arXiv on 7 August 2026 built a complete census of 4,776 operating restaurants, cafes and bars in greater Canggu and greater Ubud, then measured how many of them ChatGPT, Claude, Gemini and Perplexity actually named. 4,087 venues were never recommended in a single run. Owning a website raised a venue's odds of entering the recommended set by 1.92 times. Average star rating showed no measurable effect at that stage at all.
The study is titled Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census. It was written by Vladimir Pitenin of Norly Research, carries the identifier arXiv:2608.07069, and runs to 31 pages and 10 figures. PPC Land reported it on 23 August 2026. The paper discloses that Norly, which sells review-management and AI-visibility tools to hospitality businesses, both funded and conducted the work. The insulation measures listed are pre-registration before confirmatory collection, a dataset frozen before the models ran, adjudicated validation gates with failures reported, and published null results on hypotheses that would have been commercially convenient.
Why the denominator is the whole story
The number that matters here is not 85.6 percent. It is 4,776. Every AI visibility report published so far scores brands against a curated list of businesses that someone already considered notable, usually tens or hundreds of names. According to the paper, a catalogue without a real-world denominator can report relative prominence, but it cannot produce a population rate, which makes the phrase "never recommended" unmeasurable. If your list contains 200 famous restaurants and the assistant names 40 of them, you learn something about the 200. You learn nothing about the other few thousand businesses in that city.
Enumerating an entire market supplies the missing denominator. That is the methodological contribution, and it is the part a local business owner anywhere should take away, because it changes what a visibility score is allowed to claim. A vendor dashboard reporting that a restaurant ranks fourth on a tracked prompt is describing a position inside a set someone else drew. It is not describing the odds of appearing at all.
How 4,776 venues were counted
The population frame was defined operationally rather than editorially. It covers every place listed on Google Places under five food-service types, restaurant, cafe, bar, coffee shop and bakery, inside fixed polygons covering greater Canggu, including Berawa, Batu Bolong and Pererenan, and greater Ubud, including Penestanan, Sayan and Pengosekan. According to the paper, Google Places was chosen because it works simultaneously as the dominant consumer discovery layer and as a principal grounding source for the systems being audited.
Enumeration used the Places Nearby Search API across an adaptive grid. The API returns at most 20 results per query, so any cell that came back with exactly 20 was recursively subdivided into four half-radius cells, down to a 130 metre floor. Without that step, dense centres truncate silently and the census undercounts exactly where venues cluster. The completed grid required 760 requests. A second pass then probed roughly 220 AI-named venues that had failed to match anything in the census against Places Text Search, recovering hotel restaurants and coworking cafes that food-type enumeration misses. 70 venues in the final registry carry that lookup provenance. Each of the 4,776 entries holds a snapshot of its Google profile taken at collection time: rating, review count, price level, hours, website and business status.
96 prompts, four systems, 2,208 runs
Four production assistants were queried through their public search-grounded interfaces: OpenAI on gpt-5.2-2025-12-11 through the Responses API with the web-search tool, Anthropic on claude-sonnet-5 with the server-side web-search tool, Google on gemini-3.5-flash with Google Search grounding, and Perplexity on sonar. Two configuration choices are disclosed rather than buried. Claude's search budget was capped at two searches per run, a cost decision taken before confirmatory collection that bound in practice, since 97.6 percent of that arm's runs used both searches. Gemini's grounding API exposes no user-location parameter, so location context was carried in the query text for every system.
The instrument crossed eight personas, digital nomad, couple on a date, business meeting host, budget backpacker, family with children, dietary-constrained diner, specialty coffee enthusiast and late-night group, with six first-person paraphrase templates and two areas. That produces 96 unique queries. The confirmatory wave ran over seven calendar days with repetitions spread across days: Perplexity 10 per query, OpenAI and Gemini five each, Claude three. The totals are 2,208 runs, 12,439 valid venue mentions, 1,407,600 venue-level exposure opportunities and 9,214 successes.
85.6 percent were never recommended by any system
Those 2,208 runs produced 9,791 venue recommendations covering 689 of the 4,776 census venues. The remaining 4,087, or 85.6 percent, were never recommended by any system in any run. The share never even mentioned in passing was 84.5 percent. Narrowing the denominator barely moves it. Among venues confirmed operational the never-recommended share is 85.8 percent, and among the 2,173 established venues carrying at least fifty Google ratings it is still 72.6 percent. Widening the frame moves it the other way: a capture-recapture analysis against an independent Foursquare enumeration, corrected for duplicates and matcher recall, estimated roughly 9,400 food entities across the broadest defensible frame, under which the never-recommended share exceeds 92 percent. Every reported rate is therefore a floor, not a ceiling.
What happens after that filter is less brutal than the discourse usually assumes. The single most-recommended venue, Seniman Coffee in Ubud, accounts for 1.9 percent of all recommendations. The pooled Gini coefficient across recommended venues is 0.668. The top five capture 7.7 percent, the top ten 13.8 percent and the top 25 28.5 percent. Being recommended at all is the scarce event. Among the venues that clear that bar, the field is comparatively open.
Entry and ranking follow two different rulebooks
The paper's central structural claim is that visibility operates at two separate margins governed by different signals. A pre-registered binomial model estimated the probability that a venue is recommended in an eligible run, with engine and persona fixed effects, standard errors clustered by venue and Benjamini-Hochberg correction across nine hypothesis terms. A separate conditional logit, restricted to the 1,855 runs that recommended at least two matched venues and covering 8,450 venue-run alternatives, estimated what predicts first position once a venue is already in the answer. The table below sets the two side by side, with odds ratios as reported.
| Signal | Entry into the answer | First position within an answer |
|---|---|---|
| Own website | 1.92 (CI 1.40 to 2.62) | 1.33 |
| Review volume | 1.64 (CI 1.37 to 1.97) | 1.30 |
| Listed price information | 1.54 (CI 1.13 to 2.10) | 1.32 |
| Web-mention volume | 1.44 (CI 1.12 to 1.85) | 0.92, null |
| Average star rating | 0.89, null (CI 0.76 to 1.03) | 1.17 (CI 1.08 to 1.28) |
The short version the paper offers is that documentation admits and rating orders. Conditional on review volume and documentation, average rating shows no detectable association with whether a system surfaces a venue at all, at an adjusted p of .135. It does predict which venue is named first. Review recency leans positive at 1.41 but misses the corrected threshold at an adjusted p of .054 and is reported as suggestive rather than established. The count of review and blog domains covering a venue is null at entry at 0.90 and positive for first position at 1.23.
One coefficient is excluded from interpretation by the author, and the reason is worth reading. Venues with listed opening hours appear less likely to be recommended, at 0.54 with an adjusted p of .009. But all 70 venues added to the census through mention-driven lookup lack the hours field, because the lookup request omitted it, and those venues are recommended nearly by construction. The variable partly encodes registry provenance rather than profile completeness. The paper reports it for transparency and treats contamination of a covariate by the auditor's own discovery path as a general hazard for this kind of work.
Foursquare presence did nothing at either margin
Presence in an open point-of-interest dataset is a widely repeated assumption in local visibility work, and part of the study was designed to test it. The answer is a plain null. Listing on Foursquare returns an odds ratio of 0.84 at entry with an adjusted p of .252. Within the 480 Foursquare-listed venues in the case-control set, the graded ladder is uniformly non-significant: data-quality rating 1.12 at p .352, tip count 1.17 at .133, popularity score 1.12 at .308. Two readings are described as compatible with the data, that assistant retrieval does not touch the dataset for venue discovery in this market, or that its signal is redundant with the web presence already measured. Neither supports treating dataset inclusion as a visibility lever.
Closed businesses were recommended 93 times
Before variant recovery, 19.7 percent of valid mentions, 2,452 of them, matched nothing in the census. Every unmatched name was classified through a web-verified taxonomy: 38.4 percent were name variants of registered venues, recovered through 142 evidence-audited mappings, 3.8 percent were venues verified as permanently closed, 2.4 percent were real venues outside the study polygons, and 54.6 percent were unverifiable long-tail singletons or variant candidates below the evidence bar. After recovery, 1,511 mentions, 12.1 percent of valid mentions, remain unresolved.
Outright invention is close to absent. A single name, appearing 10 times, survived conservative scrutiny as likely fabricated, which is 0.08 percent of all valid mentions. Against that, the four systems recommended permanently closed venues 93 times across 14 confirmed-closed establishments. The paper's phrasing is that "staleness, not hallucination, is the practical failure mode." The generation layer is faithful to what retrieval hands it. The sources themselves are out of date. A closed cafe's reviews, listicle entries and blog mentions persist, and those are the same signals the entry model rewards, so the documentation trail that creates visibility keeps it alive after the business has gone. An audit scored against a curated brand list would count most of those recommendations as successes.
Ask the same question twice and get a different answer
Identical queries repeated on the same engine return substantially different venue sets. Mean top-20 Jaccard overlap runs 0.45 for Gemini, 0.40 for Perplexity, 0.29 for ChatGPT and 0.22 for Claude. Meaning-preserving paraphrases perturb results further for every engine, sharply for Perplexity at 0.19 against 0.40 and Gemini at 0.30 against 0.45. A pre-registered test-retest holdout of 16 queries across four engines, re-run two weeks later for 144 runs, found cross-period similarity comparable to the same-period rerun baseline, pooling at 0.375. The churn is sampling noise rather than drift over time, which supports treating visibility as a persistent property of a venue observed through a noisy channel.
Cross-engine agreement is low as well. Pairwise top-20 Jaccard ranges from 0.33 between Perplexity and OpenAI to 0.54 between OpenAI and Claude. Only eight venues appear in all four engines' top-20 lists, while 15 appear in exactly one. In bulk, Perplexity recommended 451 distinct venues at 6.4 mentions per run, Claude 342 at 5.8, ChatGPT 328 at 4.7 and Gemini 291 at 4.9. Census match rates ran 91 percent for Gemini, 89 percent for Perplexity, 85 percent for ChatGPT and 84 percent for Claude. The implication the paper draws is blunt: a single-shot, single-engine visibility check, which is the standard commercial product, is measuring noise.
What the assistants actually read
The 26,993 grounding citations attached to wave runs span 986 domains. Pooled across engines, the most-cited domain is not a platform but one venue's own website, finnsbeachclub.com at 5.3 percent, ahead of TripAdvisor at 2.3 percent. Regional listicle and travel-blog domains follow, thehoneycombers.com at 3.9 percent and wanderlog.com at 2.2 percent, with Reddit and YouTube at 1.3 percent each. Source diets diverge by engine: ChatGPT cites reddit.com most at 4.0 percent, Claude leads with tripadvisor.com at 8.2 percent, Gemini with wanderlog.com at 3.4 percent and Perplexity with finnsbeachclub.com at 8.1 percent. A measurement caveat applies to Gemini, whose API exposes only grounding-redirect domains, so its true sources were recovered from citation titles by deterministic mapping with unrecoverable citations dropped.
What the study does not show
The scope limits are real and the paper states them. The estimates are associations under controls, not causal effects, and established venues both maintain websites and accumulate digital footprint, so unmeasured scale could contribute to the website coefficient. The census covers the Google-listed food-service market, so unlisted micro-warungs and atypically typed outlets are underrepresented, which biases the invisibility rate downward. The audit route was the search-grounded API rather than the consumer app, and provider-side serving configurations can change without notice. The setting is two adjacent Balinese submarkets, English-language queries and a tourist and remote-worker demand profile. Replication in a structurally different market is described as required before generalising. Three pre-registered predictors, review-text similarity to query intent, cross-platform rating consistency and social presence, were never operationalised, which the paper discloses rather than works around. There is also an arithmetic inconsistency between the narrative and one figure caption on the unmatched-mention breakdown, where the caption's class counts sum to 2,860 rather than 2,452; the confirmed-closed count of 93 is identical in both places.
This is Bali, not Thailand. It is two submarkets, not a country. It is food service, not all local business. It is one snapshot in one week, with one arm capped at two searches per run by a disclosed cost decision. There is no Thai percentage in this paper and none can be derived from it.
What this means for Thai marketers
The following is Relevant Audience's reasoning, not a finding of the paper. Canggu and Ubud are tourist-heavy, English-query, independently owned food and beverage submarkets with a high density of small operators. Districts like Nimman in Chiang Mai, Thong Lo in Bangkok and Bang Tao in Phuket share that shape. That resemblance is a reason to take the mechanism seriously, and it is not a reason to assume any particular Thai number. What generalises is not the 85.6 percent. It is the structure the paper found underneath it.
Three things in that structure are checkable without any tooling. First, the largest measured entry factor was a venue owning a website, and the most-cited domain across 26,993 citations was a venue's own site. A Thai restaurant operating only through a Facebook page and a Google Business Profile has no owned document for a grounded assistant to retrieve, and that is a fixable condition rather than a permanent one. Second, average rating did not predict whether a venue appeared at all, which means a 4.8-star venue with 30 reviews and no website is competing on the wrong variable. Third, single-run checks are noise, so treating one ChatGPT answer as evidence of visibility, in either direction, is a mistake. Work on the documentation layer described on Relevant Audience's local SEO service page, and treat AI answer surfaces as a measurement problem in their own right, which is what the generative engine optimisation work covers.
What to check on your own listing
- Whether the business has an owned website at all, and whether the Google Business Profile links to it.
- Whether the profile carries price information and current opening hours, both of which are structured fields rather than marketing copy.
- Whether closed or renamed locations still exist as live profiles and stale listicle entries, because the study shows those keep getting recommended.
- Whether the same prompt run several times on several assistants returns the business, rather than judging the question on one answer.
- Whether third-party coverage exists on regional listicle and review sites, since those were the dominant non-venue citation sources.
FAQ
Does this mean 85 percent of Thai restaurants are invisible to AI?
No, and the study does not support that claim. The 85.6 percent figure describes 4,776 venues in greater Canggu and greater Ubud in Bali, measured in one seven-day wave in 2026. The paper explicitly calls for replication in a structurally different market before generalising, and no Thai equivalent has been published.
If star rating does not matter, should restaurants stop collecting reviews?
No, because the study separates rating from review volume and only rating was null. Review volume carried an odds ratio of 1.64 per standard deviation of log review count at entry and 1.30 for first position, and average star rating predicted first position at 1.17 even though it did not predict entry. The finding is about which lever does which job, not about reviews being worthless.
Is a website really more useful than a Google Business Profile?
The study does not frame it as a choice between the two, and both were present in the model. Every venue in the census came from Google Places, so a profile was the baseline condition for all 4,776, and owning a website on top of that carried the largest entry effect at an odds ratio of 1.92. The paper also cautions that this is an association under controls, not a causal effect, since larger operations tend to have both websites and wider digital footprints.
Why did the assistants recommend businesses that had closed?
Because the underlying sources were out of date, not because the models invented the venues. The paper attributes 93 recommendations across 14 confirmed-closed establishments to stale retrieval and puts outright fabrication at 0.08 percent of valid mentions, a single name appearing 10 times.
How often should AI visibility be checked to get a reliable reading?
The paper does not prescribe a frequency, but it does show why one check is not enough. Repeating an identical query on the same engine returned mean top-20 overlap of 0.22 to 0.45 depending on the system, so a single run on a single assistant sits inside that noise band and cannot distinguish a real change from resampling.







