TL;DR
- An Empirank audit published on 18 August 2026 tested 1,257 factual claims across 182 businesses that a search-enabled OpenAI configuration recommended on local queries.
- Twenty businesses, or 11.0%, carried at least one claim the evidence contradicted, and five carried at least one high-severity error.
- At claim level 24 claims (1.9%) were incorrect, 749 (59.6%) confirmed correct, 479 (38.1%) unsupported, 5 (0.4%) ambiguous and none outdated.
- The study refuses to count unsupported as incorrect because that would give an inaccuracy rate above 40% and hide the difference between a brand needing a correction and one needing better public evidence.
- Identity claims were 99.5% correct across 182 tested, while accreditation claims were confirmed only 19.2% of the time across 26, mostly for want of accessible evidence.
Twenty of 182 businesses recommended by a search-enabled OpenAI configuration carried at least one factual claim that the reviewed evidence contradicted, according to an Empirank audit published on 18 August 2026 and reported by PPC Land on 24 August 2026. That is 11.0% of the businesses tested. At the level of individual claims, the same dataset produces a very different headline: 24 of 1,257 testable claims were incorrect, a rate of 1.9%.
Both numbers are true and both describe the same audit. Which one you quote depends entirely on the denominator you pick, and that choice is the actual subject of the study. The audit, filed under the identifier AIV-009 and written by Empirank founder James Tandy, argues that a claim-level rate and a business-level rate answer different questions, and that anyone managing named clients is exposed to the second one.
What the audit counted that visibility tools do not
Most measurement of brand presence inside AI assistants counts appearances. A tool records whether a brand was named in an answer, sometimes whether it was cited, and reports a share of voice. Empirank counted something further along: whether the sentences attached to those appearances were factually true.
The distinction has a practical edge. A brand can be named, cited, and still described inaccurately in the same answer, and none of the three states are visible to a monitor that only records presence. PPC Land notes that an IAB measurement framework published on 3 August 2026 found that only 16% of brands systematically track AI visibility at all, while more than 20 vendors sell tools that disagree with one another about the same brand. Cloudflare released a dashboard on 6 August 2026 that separates how often assistants name a brand from how often they cite it. The Empirank audit adds a third axis to that pair.
How the sample narrowed
Empirank started from a pool of 3,000 sampled businesses and kept only those the model actually recommended in response to local queries. Nine recommendations were then dropped because the underlying business identity could not be resolved safely, a step taken before any accuracy judgment so that a claim is always tied to a definite entity rather than to a name that might match two businesses.
The surviving 182 businesses produced 1,515 statements. Of those, 1,257 were factual claims that could be tested against evidence. The remaining 258 were descriptive or evaluative language that no evidence set can confirm or deny, the sort of wording that calls a place popular or welcoming. Discarding that layer before counting is what keeps the error rate from being padded on either side.
Each testable claim then received one of four labels. Confirmed correct meant the available evidence directly supported it. Incorrect meant the evidence contradicted it. Unsupported meant the evidence neither confirmed nor contradicted it. Ambiguous meant the wording or the evidence did not allow a reliable decision. A fifth possible label, outdated, was applied to no claim in the set.
The distribution, and the number that carries the argument
The table below is the full claim-level result across the 1,257 testable claims, as published in the audit.
| Label | Claims | Share of 1,257 |
|---|---|---|
| Confirmed correct | 749 | 59.6% |
| Unsupported | 479 | 38.1% |
| Incorrect | 24 | 1.9% |
| Ambiguous | 5 | 0.4% |
| Outdated | 0 | 0% |
Restricting the comparison to the 773 claims where the evidence pointed decisively one way or the other changes the picture again. Within that subset, 96.9% were correct. The study reports that figure but refuses to let it stand in for the whole result, because the subset excludes by construction the 479 claims the evidence could not settle.
Why the study will not fold unsupported into incorrect
This is the part worth reading twice. Adding the 479 unsupported claims to the 24 incorrect ones would produce a single inaccuracy figure above 40%, and Empirank argues that this would be both wrong and useless. Wrong, because it treats an absence of evidence as evidence of error. Useless, because it collapses two conditions that call for completely different responses.
An incorrect claim means the model said something the evidence contradicts, and the fix runs through correcting the record. An unsupported claim means the model said something about a business that no accessible source confirms or denies. In that second case the model may well be right. What is missing is public, checkable evidence that anyone, human or machine, could use to verify the statement. Tandy restated the point in an email sent to PPC Land on 24 August 2026, writing that "unsupported does not mean false" and describing the 479 unresolved claims as a research limit rather than an accusation against the model.
For a marketer this is the genuinely useful idea in the whole study, and it inverts the instinct. Roughly four in ten claims about the audited businesses could not be checked, and the reason sits with the businesses at least as often as with the model. If a search-capable model asserts something about your company and nothing on the open web confirms it, that is a hole in your own published evidence about yourself.
Where the errors and the gaps clustered
Claim categories did not behave alike. Identity claims proved the most reliable at 99.5% confirmed correct across 182 tested claims. Accreditation claims sat at the far end, with only 19.2% confirmed across 26 claims.
The audit states the caveat on that second figure plainly: the low confirmation rate came mostly from missing evidence rather than from accreditations being disproved. A certification that no accessible page documents is not one the audit can call false. It is one the audit cannot call anything at all.
That pattern has a commercial edge in any category where credentials do real work in a buying decision. Trade licences, professional memberships, industry awards and certifications are exactly what a prospective customer weighs, and exactly the claims that tend to live on a membership registry, an association directory, or nowhere public at all. A model asked to summarise a business will often reach for those claims precisely because they sound decisive, and the audit found they are the hardest ones to substantiate.
Citations did not settle accuracy either
Empirank tested whether the presence of an attributable citation predicted whether a claim was correct. Citation-supported recommendations reached a 59.9% confirmed-correct share across 1,199 claims and 170 businesses. Recommendations without an attributable citation reached 53.4% across 58 claims and 12 businesses.
The study describes that 6.5 percentage point difference as descriptive rather than causal, and gives two reasons for the restraint. The uncited group contained only 12 businesses, too few to carry a reliable comparison. And citation presence was measured at the recommendation level, so a single cited page attached to a recommendation does not establish that every sentence inside it traces back to that page.
A second split looked at source type. Claims supported only by first-party material reached 75.2% confirmed correct across 105 claims, against 57.5% across 1,064 claims supported only by third-party material. The study flags that the two groups differ in both size and content and stops short of treating the result as proof that owned sources produce more accurate answers. That restraint is worth noticing, because the opposite conclusion would be the convenient one for anyone selling optimisation services.
What a business can actually check about itself
The audit proposes an order of operations rather than a tool, and it is checkable without any special software. Identity and location facts come first: business name, address, phone number, opening hours, service areas and branch details. Commercial proof follows, aligning services, credentials, awards and other commercial claims across the company website and authoritative third-party profiles. High-severity errors get escalated, and the same prompts are repeated after source corrections to observe whether the answer changes.
The prioritisation logic is stated in terms of customer behaviour. Errors involving location, availability, credentials, service scope and branches are more likely to change what a customer does than a descriptive flourish is. Translated into a task list, the question to ask about your own business is narrow: for each fact you would want an assistant to state, is that fact published somewhere public and machine-readable, and does it match everywhere it appears? An accreditation named only on a PDF certificate in a drawer, or an opening-hours change made on one profile and not another, is exactly the kind of input that produces an unsupported or contradicted claim.
This is where generative engine optimisation work overlaps with ordinary local SEO hygiene rather than replacing it. Consistent name, address and phone data, an accurate service list, and credentials documented on a page a crawler can reach are the same inputs both disciplines rely on. Whether that sequence actually improves answer accuracy is not something this study establishes: it measured one frozen response set once, and the retest step it describes is a proposal, not a reported result.
The limits the study puts on itself
Empirank lists its constraints rather than burying them, and they are substantial enough that the numbers should not be quoted as a general accuracy rate for AI assistants.
The response set was frozen, captured from one search-enabled OpenAI configuration in August 2026, and the study notes results may shift across models, interfaces or time. The audit covered only the 182 identity-validated businesses the model recommended, not the full 3,000-business pool. Citation presence was assessed at the recommendation level. The uncited comparison group held just 12 businesses. Quality control ran through independent AI agents rather than human reviewers, a choice the study discloses in its own limitations section, which means an audit of machine output was itself reviewed by machines.
The study also does not name the markets or languages its local queries covered, and nothing in the published material addresses Thailand or any other specific country. Instability elsewhere in the field compounds the caution: SISTRIX measurement covered in May 2026 found ChatGPT rotating 74% of its citations per week, against 56% for Google AI Mode, which makes any frozen snapshot a photograph rather than a rate.
Why an aggregate rate and a client-level rate feel so different
Agencies and in-house teams do not manage claim totals. They manage named businesses, and one wrong sentence attached to one of them is a problem no matter how small it looks inside a pooled percentage. A claim-level error rate of 1.9% is a property of a dataset, while a client roster is what a team is answerable for. That is the structural point the audit presses, and the same arithmetic shows up across other published work in the field.
The stakes are not purely editorial, because recommendations move traffic. Similarweb clickstream research published in June 2026 found brands recommended by ChatGPT were 2.5 times more likely to receive a visit within seven days, with 55.9% of that AI-influenced traffic arriving through branded search rather than a referral link. Consumer verification behaviour then decides what happens next: Yelp research with Morning Consult published in April 2026 found only 15% of United States adults trust AI search results a lot, even though 65% had used such a tool in the previous six months and 63% double-check results against other sources. An incorrect claim inside a recommendation does not simply sit unread. It gets checked.
What this means for Thai marketers
The audit does not cover Thailand, does not say which markets its local queries came from, and should not be presented in a Thai pitch as a measured Thai error rate. What travels is the method, not the number.
The part that transfers most directly is the unsupported category, because the conditions that produce it are common in the Thai market. Business facts here are often split across a Facebook page, a LINE Official Account, a Google Business Profile and a website, with hours and branch lists updated in one place and not the others. Credentials issued by Thai regulators, professional bodies or industry associations frequently exist only in a Thai-language registry, a PDF, or a certificate on a wall, which puts them in exactly the state the audit describes as unsupported rather than false.
There is a bilingual dimension too. If a business publishes its service list, hours and credentials in Thai only, an assistant answering an English-language query has thinner ground to stand on, and the reverse holds for English-only sites serving Thai customers. The audit does not test this, and no published figure quantifies it, so treat it as a reason to check your own evidence rather than as a finding. The check itself costs nothing: read what your public profiles say about you, and ask whether a stranger could verify each claim from those sources alone.
FAQ
Does this mean OpenAI gets 11% of businesses wrong?
No. Empirank found that 11.0% of the 182 audited businesses carried at least one claim the evidence contradicted, which is a business-level exposure rate, not a per-answer error rate. At claim level the incorrect share was 1.9%, or 24 of 1,257 testable claims, and both figures come from the same frozen dataset captured from one search-enabled OpenAI configuration in August 2026.
What does an unsupported claim actually mean?
It means the reviewed evidence neither confirmed nor contradicted the statement, which is not the same as the statement being false. Empirank recorded 479 unsupported claims, 38.1% of the total, and Tandy told PPC Land on 24 August 2026 that "unsupported does not mean false" and that these represent a research limit rather than an accusation. In practice an unsupported claim about your business often points to missing public evidence about you rather than to a model error.
Why were accreditation claims only 19.2% confirmed?
Mostly because the evidence was missing rather than because the accreditations were disproved, which the audit states directly. Only 26 accreditation claims were tested, a small base, and certifications tend to be documented on membership registries, association directories or nowhere publicly accessible at all.
Does the study say what to do about it?
It proposes an order of operations without claiming to have proved it works: verify identity and location facts first, then align services, credentials, awards and commercial claims across the company site and authoritative third-party profiles, escalate high-severity errors, and repeat the same prompts after corrections. The study measured one frozen response set once, so the retest step is a proposal rather than a reported result.
Do these findings apply to Thailand?
The study does not say. It does not name the markets or languages behind its local queries, does not mention Thailand, and covers 182 businesses drawn from a 3,000-business pool under one model configuration, so no Thai-specific rate can be read out of it.







