TL;DR
- Pew Research Center published an analysis on 20 August 2026 finding signs of AI authorship or editing on about 10% of 490,000 English-language webpages sampled from Common Crawl in July 2026.
- Among pages carrying a publication date after ChatGPT's November 2022 release, the share rose to more than one-third.
- By domain, the January 2026 reading was 9.35% on .com, 4.59% on .org, 1.03% on .edu and 0.76% on .gov, against roughly 1% or below for all four before ChatGPT.
- Pew's threshold catches substantial AI editing as well as full AI generation, so a human-written page polished in Google Docs or Word falls in the same category.
- Graphite put primarily AI-generated new English articles at 49.9% for Q1 2026 and an unreviewed preprint found 35% of new websites AI-assisted by mid-2025, all leaning on the same detector family.
Pew Research Center found signs of AI authorship or editing on about 10% of the English-language webpages it sampled in July 2026, and on more than one-third of the pages carrying a publication date after ChatGPT was released. The analysis published by Pew Research Center came from its Data Labs team on 20 August 2026, and it ran almost half a million pages from the Common Crawl web archive through an open-weight AI detection model. Pew states directly that the finding describes rates across large collections of text and cannot label any single document.
That caveat is the part most likely to fall out of a summary, so it belongs at the top. The 10% is a population estimate. It is not a verdict on a page you published, and it does not separate text a model wrote from text a person wrote and then cleaned up with an AI editing feature. Both land in the same bucket.
What Pew actually measured
Pew analysed roughly 490,000 English-language webpage texts drawn from the Common Crawl web repository, using snapshots collected between January 2021 and July 2026. Beginning the series almost two years before ChatGPT launched in November 2022 gives the work a baseline. Whatever the model flags in the 2021 samples is the floor: the rate at which ordinary human writing happens to look, statistically, like something that did not yet exist.
The tool was the Open Pangram detector, an open-weight model created by Pangram. Pew describes it as looking for statistical differences in how human authors and AI models use language, including word choice and sentence structure, rather than running down a list of banned words. Pew also says outright that detection models sometimes misclassify human-written documents as AI-influenced and the reverse. That is the reason the essay reports shares of very large samples and never makes a call on an individual page.
Two headline numbers came out of the July 2026 snapshot. Across all sampled pages, 10% showed significant signs of AI authorship. Restricted to pages carrying a publication date after ChatGPT's release, the share rose to more than one-third. The distance between those two figures is arithmetic rather than a discovery. A random slice of the web holds a great deal of material written before a generative model could have touched it, and that older material pulls the overall rate down.
The concentration sits on the commercial web
Pew broke the detection rate down by top-level domain and plotted the results as six-month averages. Before ChatGPT, the four main domain types sat close together, all at or near 1%. By the reading plotted at January 2026 they had separated, with .com climbing furthest while .edu and .gov barely moved. The table below sets the pre-launch reading against the most recent point in Pew's published series.
| Top-level domain | Reading plotted at July 2022 | Reading plotted at January 2026 |
|---|---|---|
| .com | 1.07% | 9.35% |
| .org | 0.63% | 4.59% |
| .edu | 0.64% | 1.03% |
| .gov | 0.25% | 0.76% |
Pew's own reading of that split is that around one in ten .com pages in the 2026 samples show signs of AI authorship, roughly double the .org share and about ten times the .edu and .gov rates. For anyone working in marketing, that is the uncomfortable half of the finding. The concentration sits precisely on the part of the web that commercial content teams publish to, which is also the part of the web most search work touches.
The linguistic tells, and what they are worth
Pew tracked how often particular markers appear across pages published after 30 November 2022, counted per 10,000 words and plotted as six-month averages. Comparing the January 2026 reading with the January 2023 reading gives the following movement.
- Em dashes rose from 5.79 per 10,000 words to 11.19, roughly twice as frequent.
- Oxford commas rose from 34.04 to 55.51, which Pew summarises as a 63% increase.
- A defined list of AI-typical vocabulary, 27 terms that include "delve", "interplay" and "testament", rose from 11.94 to 26.02, more than doubling.
- Negative parallelism, the "it's not just X, it's Y" construction, rose from 0.87 to 2.36, close to tripling while remaining rare in absolute terms.
Pew is blunt about the limits here. Human writers use em dashes, Oxford commas and lists of three, and no single marker settles the authorship of a document. The claim is about averages across large bodies of text. Pew also notes that the detector it used has more to work with than a handful of surface signals, since it learns subtler statistical patterns in word choice and sentence construction.
There is a practical trap in the marker list that is worth naming. A rate that rises across the whole web says nothing about the page in front of you, and a writer who removes every em dash from a document has changed a statistic by one observation, not changed whether the document is accurate or worth reading. Pew did not examine whether any individual marker, negative parallelism included, has become more or less useful for identification as these features spread, and the essay says so.
AI editing counts the same as AI generation
This is the second structural limit, and for a working content team it matters more than the headline. Pew's threshold catches substantial AI editing as well as full AI generation. AI editing is now a built-in feature of Google Docs and Microsoft Word, which means a page a person researched, drafted and argued, then ran through a grammar and rewrite tool, can land in the same category as a page produced from a single prompt.
None of the published estimates separate the two. That is not an oversight the studies hide; Pew's framing is about pages "written or substantially edited by AI", and the wording is deliberate. It does mean any sentence of the form "10% of the web is AI-written" overstates what was measured. The honest version is that about 10% of sampled pages carry enough statistical fingerprints of model-shaped language for a detector to flag them, from whatever mixture of drafting and editing produced that.
Three estimates that do not agree
Pew's figure sits alongside two other recent attempts at the same question, and the three do not line up. As reported by Search Engine Journal on 24 August 2026, Graphite, an SEO firm, estimated that 49.9% of newly published English-language articles in the first quarter of 2026 were primarily AI-generated. A preprint from Imperial College London, the Internet Archive and Stanford, which has not been peer reviewed, found that 35% of newly published websites were detected as AI-generated or AI-assisted by the middle of 2025.
Some of the spread is definitional. Pew sampled webpages from a general crawl and reported both a whole-sample rate and a post-ChatGPT-pages rate. Graphite measured newly published articles. The preprint measured newly published websites. Those are three different populations, and ranking them against each other is not a like-for-like comparison.
The rest of the spread is harder to explain away, and there is a dependency running underneath all three. Every one of these estimates leans on Pangram in some form, with Graphite also using Copyleaks and GPTZero. When several studies share a detector family, their errors are correlated rather than independent, so agreement between them is weaker evidence than it looks and disagreement is harder to arbitrate. Nobody has published a ground-truth sample of the open web that would let a reader decide which number is closest.
What this does and does not imply for a content process
Start with what it does not imply. Pew did not test how flagged pages perform in search. It did not test whether readers can tell the difference, whether flagged pages convert worse, or whether any platform treats a detection score as a signal. Any claim that this analysis shows AI content being penalised is an addition to the source, not a reading of it.
What it does imply is narrower and more useful. The statistical signature of model-shaped prose is now common enough on commercial domains that it no longer distinguishes anything. If a page's only claim on a reader's attention is that it is fluent, competent and organised, it is now competing against a very large volume of fluent, competent, organised pages. The scarce inputs are the ones a model cannot supply on its own: a named source with a date attached, a number someone actually verified, an original observation, a document read in full rather than summarised second hand.
That reframes the editorial question from "does this read as human" to "does this contain anything a reader could not get from a generic answer". A useful content marketing process handles that at the brief stage, by deciding what specific evidence a piece has to carry before anyone writes a word, rather than at the copy-editing stage by hunting punctuation. Removing markers from a thin page leaves a thin page.
What Google's published position says
Google's guidance on AI-generated content in Search states that its focus is on the quality of content rather than how it was produced, and that using automation to generate content with the primary purpose of manipulating search rankings falls under its spam policies. Both halves are worth holding together, because they say a detection rate is not the operative test in either direction.
Neither Pew nor the trade coverage of the analysis made any claim about rankings or penalties, and nothing in the data speaks to search performance. Pew's closing observation is simpler than that: whether a page is accurate, useful and worth publishing is still settled by reading it.
What this means for Thai marketers
The first thing to be clear about is scope. Pew sampled English-language webpage texts only, so none of these rates describe Thai-language content, and no equivalent published rate for Thai exists in this analysis. Anyone quoting the 10% figure in a Thai-market deck should say which language it refers to.
The implication is still real for two reasons. Bilingual sites publish an English half, and that half sits inside exactly the population Pew measured, on exactly the .com domains where the rate climbed fastest. And Thai teams working on SEO in Thailand often produce Thai content by drafting or translating with the same tools, which means the underlying production pattern is shared even where the measurement is not.
The second implication is about local evidence. If generic, competent prose is now abundant in English, the material that stays scarce in a Thai-market page is the local specificity: a Thai price, a Thai regulation with its issuing body named, a Bangkok delivery constraint, a platform behaviour that differs here. That is content a model has no reliable way to invent, and it is the part of a page a Thai reader can immediately test against what they already know.
What the analysis did not say
Pew did not say that 10% of webpages were written entirely by AI, that any named site or publisher uses AI, that flagged pages are lower quality, that the detector is accurate on individual documents, or that any search engine measures or acts on these markers. It did not measure non-English content, and it did not resolve the disagreement between its own figure and the two competing estimates.
FAQ
Does a page with em dashes get flagged as AI?
No. Pew says explicitly that no single marker identifies a document, because human writers use em dashes, Oxford commas and lists of three all the time. The em dash rate Pew published, 5.79 per 10,000 words in early 2023 rising to 11.19 in early 2026, is an average across hundreds of thousands of pages, not a threshold applied to one article.
Is 10% of the internet written by AI?
Not as stated. Pew found signs of AI authorship or editing on about 10% of the English-language pages it sampled in July 2026, which includes pages a person wrote and then substantially edited with an AI tool. The share rises to more than one-third when the sample is limited to pages published after ChatGPT's release, and it says nothing about non-English pages.
Why do Graphite and Pew report such different numbers?
Because they measured different populations with overlapping tools. Graphite put primarily AI-generated newly published English-language articles at 49.9% for the first quarter of 2026, while Pew reported a rate across a general sample of webpages. All the current estimates rely on Pangram in some form, so their errors are related rather than independent, and no published ground-truth sample exists to settle which is closest.
Will Google penalise my pages if a detector flags them?
Pew did not address rankings at all, and the analysis contains no data on search performance. Google's published guidance on AI-generated content says it judges pages on quality rather than on how they were produced, and treats automation aimed at manipulating rankings as a spam policy matter, which is a different test from any detector score.
Does this apply to Thai-language content?
Not directly. Pew analysed English-language webpage texts only, so there is no Thai rate in this data and the study does not say what a Thai-language sample would show. The domain pattern is still relevant to Thai teams because the .com sites they publish English pages on are where the measured rate rose fastest.







