New York passes a Stealth Crawler Prohibition Act as unidentified bots ignore robots.txt

New York passes a Stealth Crawler Prohibition Act as unidentified bots ignore robots.txt

SEOAugust 5, 2026
By Antonio Fernandez

New York State passed a Stealth Crawler Prohibition Act in June 2026, and the bill is now waiting on Governor Kathy Hochul's signature. Digiday reported the detail on 4 August 2026: the act would require software that retrieves, scrapes or accesses a website, including AI agents, to disclose its identity and its purpose, with civil penalties of up to $15,000 per day for each violation, enforceable by the New York attorney general's office.

The bill is not law yet: Digiday did not report a signature or an effective date.

What a stealth crawler is

Digiday, in its explainer published on 4 August 2026, defines a stealth crawler as a web crawler that scrapes content from sites without identifying itself or its purpose. These bots bypass or ignore robots.txt, and they arrive looking like ordinary human visitors rather than announcing themselves as automated systems.

That undercuts most of the advice site owners have been given about AI crawling. A robots.txt directive is a request addressed to a crawler that has already chosen to identify itself, and nothing in the file binds a bot presenting as a browser.

What the New York act would require

Digiday reported on 4 August 2026 that the New York Stealth Crawler Prohibition Act passed in June 2026 and defines stealth crawlers broadly, covering any software that retrieves, scrapes or accesses a website, including AI agents. The obligation it creates is disclosure, the penalty is civil at up to $15,000 per day for each violation, and enforcement sits with the New York attorney general's office.

The source is silent on two things a site owner would want to know: when the act would take effect if signed, and whether any enforcement mechanism reaches crawler operators based outside the United States.

A federal bill exists but has not moved

Digiday also reported that a Stealth Bot Prohibition Act was introduced in the US House in July 2026 and is still seeking sponsors. That is the extent of what the source says: no committee progress, no timetable, no comparison with the New York text.

The traffic figures come from companies that sell bot mitigation

Cloudflare puts bot traffic at more than half of all web traffic. Human Security, a cybersecurity firm, reported that AI scraper traffic grew 597% from January to December 2025, that AI-driven traffic overall grew 187% in 2025, and that scraping attacks affect nearly 20% of the median organisation's site traffic.

Both companies sell bot mitigation, so both have a commercial interest in a large bot-traffic number. That does not make the figures wrong, but no independent measurement is on the table here, and Digiday does not report how any of them were measured or over what sample.

Here is what has actually been reported, separated from who reported it:

The traffic figures come from companies that sell bot mitigation
ItemStatus or sourceWhat is reported
New York Stealth Crawler Prohibition ActPassed June 2026, awaiting Governor Hochul's signatureBots, including AI agents, must disclose identity and purpose; up to $15,000 per day for each violation; New York attorney general enforces
Stealth Bot Prohibition Act (federal)Introduced in the US House, July 2026Still seeking sponsors; no further detail reported
Share of web traffic that is bot-basedCloudflare, which sells bot mitigationMore than half of all web traffic
Growth in AI scraper trafficHuman Security, which sells bot mitigation597% from January to December 2025
Growth in AI-driven traffic overallHuman Security187% during 2025
Scraping attacks as a share of site trafficHuman SecurityNearly 20% for the median organisation

What this means for Thai marketers

Digiday does not address Thailand or Thai law anywhere in the piece, so this section is reasoning rather than reporting. A New York statute governs conduct New York enforcement can reach, and a Thai-facing site scraped by an unidentified bot gains no remedy from it. Any indirect benefit would arrive slowly, if large crawler operators changed global behaviour to stay compliant in one big market.

The useful takeaway applies today. If the plan for protecting content is a block list in robots.txt, that plan covers the polite crawlers and misses the ones this legislation was written for. Which AI systems you actually want reading your content is a strategy question, and it belongs alongside the rest of your AI search work rather than in a file nobody revisits.

There is a second cost that gets ignored: undisclosed crawler traffic is server load you pay for, and sessions that distort your analytics baselines. If organic numbers move in ways on-page work does not explain, check bot filtering before rewriting anything, work that sits close to SEO and generative engine optimisation diagnostics.

Frequently asked questions

Has the New York bill become law?

No. Digiday reported on 4 August 2026 that the act passed in June 2026 and is awaiting Governor Kathy Hochul's signature. The source does not state when a signature is expected, or when the law would take effect if signed.

Does this protect a Thai website?

The source says nothing about Thailand or Thai law, so no protection can be claimed from it. This is a New York state statute enforced by a New York state authority, and any effect on Thai-facing content would be indirect.

Will blocking AI crawlers in robots.txt stop stealth crawlers?

No, and that is the definitional point of the story. Digiday describes stealth crawlers as bots that bypass or ignore robots.txt and present as ordinary human visitors, so a file addressed to self-identifying crawlers has no effect on them.

Are the bot-traffic statistics independently verified?

No. Every figure in the piece is attributed to Cloudflare or to Human Security, and both sell bot mitigation products. Digiday does not report the measurement method, sample or geographic scope behind any of them.

Is there anything I need to do today?

Nothing is required of a site owner by a bill that has not been signed. The practical step, and this is reasoning rather than a source recommendation, is to stop treating robots.txt as a content-protection layer and to check server logs and analytics for unidentified traffic you already pay to serve.

The short version

One US state has moved first on forcing bots to say who they are, a federal version is stalled at the sponsor-hunting stage, and the numbers describing the problem come from vendors selling the fix. Read the full report from Digiday for the legislative detail. If you want a clear picture of which crawlers reach your site and what your content does inside AI answers, that is worth settling before the next round of content spend.

Antonio Fernandez

Antonio Fernandez

Founder and CEO of Relevant Audience. With over 15 years of experience in digital marketing strategy, he leads teams across southeast Asia in delivering exceptional results for clients through performance-focused digital solutions.

Share to:
Copy link: