Research finds ChatGPT's own search index gives licensed publishers no advantage

Research finds ChatGPT's own search index gives licensed publishers no advantage

geoAugust 12, 2026
By Antonio Fernandez

Search Engine Journal reported on 12 August 2026 that research from the French SEO consultancy Resoneo found no difference in how OpenAI's own search index serves licensed publishers compared with sites that hold no licensing deal at all. Resoneo reverse-engineered ChatGPT's network traffic, using server tags to work out which retrieval pipeline had fetched each search result, and reported that format, length and freshness looked the same whether the page came from a small Italian publisher or from a large outlet such as Reuters or The Guardian.

That finding moves ChatGPT visibility out of the business development column and into the on-page column. If OpenAI's index is not sorting sites by whether their owner signed a contract, the question worth asking is what it does keep about a page. Resoneo's answer is that it keeps very little: roughly 200 characters per page, built from elements a marketing team already controls.

What Resoneo measured, and on what sample

Resoneo built its dataset by capturing ChatGPT sessions during July 2026 and inspecting the network traffic sitting behind the answers, according to the Search Engine Journal report published on 12 August 2026. Server tags inside that traffic identified which retrieval system had produced each individual search result, and that is what let the team separate OpenAI's in-house index from results scraped out of Google.

The sample was 1,249 ChatGPT answers captured in July, 534 pages that ChatGPT cited in its answers, and 16,407 search results pulled from paid accounts running in thinking mode. A further 463 pages carrying H1 headings formed the basis of the snippet analysis. Tests ran across free accounts and paid accounts, in more than one country, and in logged-out sessions. The logged-out condition matters, because it takes personalisation off the table as an explanation for which pages appeared.

What Resoneo did not have was any statement from OpenAI. The whole picture was assembled from the outside, by watching traffic, and every label in it is an internal tag whose meaning OpenAI has not confirmed. Read the study as a careful reconstruction rather than as documentation.

The licensing finding: one pipeline, licensed and unlicensed alike

The central result Search Engine Journal reported on 12 August 2026 is that OpenAI's in-house search index showed no difference between licensed and unlicensed sites in format, length or freshness. Small Italian publishers surfaced through the same pipeline that carried Reuters and The Guardian. There was no separate lane for the publishers OpenAI has commercial arrangements with, at least not one visible in how results were served.

For the last two years the working assumption in a lot of marketing rooms has run the other way. Every announced content deal between an AI company and a publisher reinforced a picture in which AI answer visibility was something bought at the corporate level, and where a mid-sized business had no realistic entry point. Resoneo's data undercuts that assumption for the specific case of OpenAI's own index. The implication drawn in the reporting is that small sites without any OpenAI agreement are already sitting inside the index that powers most free-user results.

Worth stating precisely what this does and does not claim. It says the retrieval layer treats the two groups the same. It does not say licensing has no commercial value to a publisher, and it does not say a small site and a national newspaper are equally likely to be cited in a finished answer. Citation depends on far more than whether a page can be retrieved.

Free accounts and paid thinking mode are running different machines

Resoneo's traffic analysis, as reported by Search Engine Journal on 12 August 2026, found that the retrieval mix changes sharply depending on which ChatGPT a person is using. On free accounts, OpenAI's in-house index, tagged internally as "labrador", handled most search results. In paid thinking mode the split inverted: roughly 75 percent of results came from Google scraping and about 24 percent from the in-house index.

The mechanism this implies is that two users can ask ChatGPT the same question and be served from substantially different retrieval systems, with different coverage and different criteria for what surfaces. Anyone who has ever pasted a screenshot of a ChatGPT answer into a client report should sit with that for a moment. A test run on a paid account in thinking mode is largely a test of what Google surfaces. A test run on a free account is largely a test of OpenAI's own index. Those are two separate findings wearing the same interface.

The practical consequence for reporting is that account tier and mode belong in the methodology line of any ChatGPT visibility check, alongside the country and the date. Without them, a before and after comparison can move purely because someone switched accounts between the two runs.

What OpenAI's index actually stores about a page

The most operationally useful part of the Resoneo research, as summarised by Search Engine Journal on 12 August 2026, is the description of what a single record in OpenAI's in-house index contains. It is small. Resoneo put it at roughly 200 characters per page, typically the page title plus the H1 heading, sometimes a publication date, and sometimes the alt text of the first image. The table below sets out each stored element against how often Resoneo found it present.

What OpenAI's index actually stores about a page
Stored elementPresence in Resoneo's sampleWhat the marketer controls
Whole index recordRoughly 200 characters per pageThe total budget the page gets to describe itself
Page title plus H1 headingH1 present on 83.6 percent of pagesBoth are editable on any CMS
Publication date11 percent of pagesExposed or hidden by the template
First image alt text9 percent of pagesWritten by whoever uploads the image

Three of the four items on that list are things an in-house team can change this week without a developer. The fourth, the publication date, usually depends on the template and is a small ticket rather than a project. There is no keyword density term in the record, no backlink count, no author bio. At the retrieval layer, a page is a title, a heading, possibly a date and possibly a short line of alt text.

How to audit your own pages against the four stored elements

Given a record that short, the audit is short too. Working through your main commercial pages and your best performing articles, check each of the following.

  1. Is there an H1 at all? Resoneo found one on 83.6 percent of the pages it examined, which means roughly one page in six had none. Design systems that style a large heading with a div, and page builders that emit two H1 elements or zero, both produce this. View the rendered source and confirm exactly one H1 exists.
  2. Does the H1 repeat the title, or add to it? If the index stores both, and the whole record is about 200 characters, a title and H1 that are word for word identical spend the budget twice on the same information. A title that carries the product or service and an H1 that carries the specific question the page answers cover more ground inside the same space.
  3. Does the title stand alone with no context? Titles written for a SERP lean on the brand name in the suffix and on the surrounding snippet. Pulled into a 200 character record with nothing around it, "Our Approach" or "Services" describes nothing. Read each title as though it is the only sentence a machine will ever see about that page, because in this pipeline it roughly is.
  4. Is a publication date exposed in the markup? Resoneo found a date on only 11 percent of pages. Many templates render a date in the visible layout but bury it in a way that is hard to extract, and many evergreen service pages carry none. If freshness is part of why your page deserves to be picked, the date has to be present to count.
  5. Does the first image carry meaningful alt text? Note the word first. Not the hero background, not the logo, whichever image the markup reaches first. The first image alt text was present in the record for 9 percent of pages. Where it exists it is free extra description; where it says "image1" or is empty, that slot is wasted.

None of this is new advice. What is new is the reason for it. These are no longer just accessibility and SEO hygiene items, they are the fields that a large answer engine keeps about the page, and there are only four of them.

Why 83.6, 11 and 9 percent are facts about the web, not ranking factors

This is the point most likely to be misread. The percentages Resoneo reported describe how often each element was present across the pages studied. They measure the state of the web, not the weight OpenAI assigns to anything.

Reading them as ranking factors produces the wrong conclusion in both directions. A team could look at 11 percent for publication dates and decide dates barely matter, when the number simply says most pages do not expose one. Or a team could see 9 percent for alt text and conclude that alt text is a strong signal because it is rare, which the data does not support either. Resoneo measured presence. Nothing in the reporting attaches a weight to any of the four elements.

What the numbers do support is an opportunity argument. If a date appears on 11 percent of pages and yours is in that group, your record carries a piece of information most competing records do not have. Whether the system rewards that is unknown.

What the research does not settle

Being straight about the limits is part of using this well.

  • It is not an OpenAI disclosure. One consultancy reverse-engineered network traffic. OpenAI has not confirmed the pipeline names, the record contents or the split between systems.
  • The tags are internal names. "labrador" is a label seen in traffic. Its meaning is inferred, and it could be renamed or retired without notice.
  • The window is a single month. The captures come from July 2026. Retrieval mixes at AI companies have changed on far shorter timescales than that, and the 75 to 24 split in thinking mode is a snapshot rather than a stable ratio.
  • Retrieval is not citation. The study covers what gets fetched and what the index holds. Which of the retrieved pages a model ends up quoting in an answer is a separate step that this work does not describe.
  • No causal test was run. Nobody added an H1 to a page and watched what happened. The relationship between the four stored elements and being picked remains untested here.

What this means for Thai marketers

The Search Engine Journal article contains no Thailand-specific data. Resoneo's tests covered more than one country and its publisher examples were Italian, and no Thai sample or Thai language finding is reported. Anything below is reasoning from the general finding, not something the source states about Thailand.

The reasoning runs like this. A Thai SME, a local publisher, a clinic, a property developer: none of them will ever sign a content licensing agreement with an AI company. Under the older assumption that AI visibility follows media deals, that closed the door. Under Resoneo's finding, the door was never the deal. The in-house index that serves most free-account traffic appears to carry small unlicensed sites on the same terms as large ones, and free accounts are where a very large share of everyday consumer usage in Thailand sits.

The four stored elements are also unusually cheap to fix on Thai language pages. Titles and H1 headings on Thai sites are frequently copied from an English template and left generic, publication dates often disappear inside theme styling, and alt text on Thai pages is empty far more often than not. Those are editorial fixes, priced in hours, and they do not require a rebuild. For a market where the gap to international competitors is usually budget rather than skill, a visibility lever that costs a copy pass is worth taking seriously. Teams working on ChatGPT SEO and broader generative engine optimisation can fold this check into a normal content review rather than treating it as a separate discipline.

Questions marketers are asking about the Resoneo findings

Does this mean I do not need a licensing deal with OpenAI to appear in ChatGPT?

For appearing in the in-house index, Resoneo's data says licensed and unlicensed sites were served identically in format, length and freshness. The reporting draws the implication that small sites without deals are already in the index that powers most free-user results. It does not follow that licensing has no value to a publisher for other reasons, and being retrievable is not the same as being quoted in an answer.

Is any of this confirmed by OpenAI?

No. Search Engine Journal reported it on 12 August 2026 as research by Resoneo based on reverse-engineering ChatGPT's network traffic. The pipeline names are internal tags seen in that traffic, and their meaning has not been confirmed by OpenAI. Treat the specific percentages as one consultancy's measurement from a single month.

Was any of this tested in Thailand or on Thai language pages?

The source did not say. Resoneo ran tests across multiple countries, but no Thai market data and no Thai language results appear in the Search Engine Journal report. The publisher examples given are Italian.

Do I have to do anything right now?

Only if your pages fail the basics. Confirm each page has exactly one H1, that the title makes sense read alone, that a publication date is present in the markup where freshness matters, and that the first image has real alt text. If those already hold across your site, this research changes nothing about your workload and simply explains why they were worth doing.

Why did my ChatGPT visibility test give a different answer to a colleague's?

Account tier and mode are the first thing to check. Resoneo found that free accounts were served mostly by OpenAI's in-house index, while paid accounts in thinking mode took roughly 75 percent of results from Google scraping and about 24 percent from the in-house index. Two people on different plans are querying different retrieval systems, so record the account type, the mode, the country and the date with every test.

The short version

A 200 character record is a brutal constraint, and it is also a generous one. It means the page has four chances to describe itself, and every one of them is editable by the people who already own the content. Resoneo's work suggests the barrier to appearing in the ChatGPT search index is not a contract signed in a boardroom, it is whether the title says something, whether the H1 exists, whether the date is visible to a parser and whether the first image was ever given a caption worth reading.

If you want that audit run properly across a site rather than spot-checked on a few pages, our team handles it as part of ongoing SEO work in Thailand. Get in touch and we will look at how your pages currently describe themselves.

Antonio Fernandez

Antonio Fernandez

Founder and CEO of Relevant Audience. With over 15 years of experience in digital marketing strategy, he leads teams across southeast Asia in delivering exceptional results for clients through performance-focused digital solutions.

Share to:
Copy link: