TL;DR
- Robots.txt controls crawling only: Google states a disallowed URL can still be indexed without a description if other sites link to it.
- Google caches robots.txt for up to 24 hours, reads at most 500 KiB, and ignores crawl-delay; a 4xx response (except 429) means no restrictions.
- Within a group, the longest matching path wins; on a tie, Google applies the least restrictive rule, so Allow beats Disallow.
- A Googlebot group replaces the * group for Googlebot instead of merging with it; AI crawlers such as GPTBot and ClaudeBot each need their own group.
A robots.txt file is a plain-text file at the root of a host that tells crawlers which URL paths they may fetch. It controls crawling, not indexing: a URL blocked in robots.txt can still show up in Google results, usually without a description, if other pages link to it. This robots.txt guide covers the syntax, how Google picks which rule wins, the user-agent names of AI crawlers, how to test the file, and the mistakes that cause the most damage.
What robots.txt does, and what it does not do
Think of robots.txt as a request about traffic, addressed to well-behaved crawlers. Before Google crawls a site, it downloads the robots.txt file and reads which parts of the site it is allowed to request. Google's crawlers follow the Robots Exclusion Protocol, which is standardized in RFC 9309. The practical uses are narrow: keeping crawlers away from endless filter combinations, internal search result pages, cart and checkout steps, and other URLs that waste server capacity without helping anyone find your content.
Three limits matter more than the syntax itself.
- It is not a way to hide pages. Google's documentation states that a page blocked by robots.txt can still be indexed if other sites link to it. The URL, and sometimes anchor text from those links, can appear in results with no description. If you want a page out of search, use a noindex rule, password protection, or remove the page.
- It cannot enforce anything. Reputable crawlers obey it. Scrapers and malicious bots can ignore it, so it is never a security control. A robots.txt file is also publicly readable, which means listing a secret folder in it advertises the folder.
- Crawlers read it differently. Googlebot understands wildcards and the Allow rule. Not every crawler does, and each interprets edge cases in its own way.
Crawling and indexing are separate steps
Crawling means fetching a URL. Indexing means deciding to store and possibly show it. A noindex directive lives inside the page, either as a robots meta tag or an X-Robots-Tag HTTP header, so a crawler has to fetch the page to see it. That creates the most common trap: you add Disallow for a section and noindex on its pages at the same time. Google can no longer fetch those pages, never reads the noindex, and may keep the bare URLs in the index. To remove a page from search, leave it crawlable until it has dropped out, then decide whether to block it.
The reverse also applies. If you block a page to save crawl effort and later decide you want it ranking, the page cannot rank on its content, because Google never reads it.
Where the file lives and how Google fetches it
The file must be named robots.txt and sit at the top level of the host, for example https://www.example.com/robots.txt. A file in a subfolder is ignored. The rules apply only to the host, protocol and port that served the file, which has several consequences:
- A file on www.example.com does not cover example.com, shop.example.com, or http://www.example.com. Each subdomain needs its own file.
- A file on a non-standard port covers only that port.
- The URL is case-sensitive, so use lowercase robots.txt.
The file must be UTF-8 plain text. Google enforces a size limit of 500 KiB and ignores anything past it, so very long files should be consolidated into broader rules. Google generally caches the file for up to 24 hours, which is why a change does not take effect the minute you upload it.
What the HTTP status code does to your rules
Google's handling of the response code is documented and worth knowing, because a broken server can change your crawl behaviour without any edit to the file.
- 2xx: the file is processed as served.
- 3xx: Google follows at least five redirect hops, then treats the result as a 404 for robots.txt. Redirects done with JavaScript or meta refresh are not followed.
- 4xx (except 429): treated as if no robots.txt exists, so there are no crawl restrictions. A 403 on the file does not block your site; it removes your rules.
- 5xx: for the first 12 hours Google stops crawling the site while it keeps retrying. After that it uses the last good version for up to 30 days. If the errors continue past 30 days, Google behaves as if there is no file when the site is otherwise available.
Robots.txt syntax: the four fields Google supports
Each line is a field name, a colon and a value. Field names are case-insensitive, but path values are case-sensitive. Comments start with #. Google supports four fields, and it does not support others such as crawl-delay.
- User-agent: names the crawler the following rules apply to. An asterisk means every crawler that has no more specific group.
- Disallow: a path the named crawler must not fetch. An empty Disallow value is ignored.
- Allow: a path the crawler may fetch, used to carve an exception out of a broader Disallow.
- Sitemap: the full, absolute URL of a sitemap. It can sit anywhere in the file, is not tied to a user agent, and can be repeated for several sitemaps.
A minimal, valid example might read: User-agent: * then Disallow: /cart/ then Disallow: /search then Sitemap: https://www.example.com/sitemap.xml. Paths must start with a slash and match from the start of the URL path, so Disallow: /search blocks /search, /search/results and /searchable-guide alike. Add a trailing slash when you mean a folder.
Wildcards
Google and the other major search engines support two special characters in paths. The asterisk matches zero or more of any character. The dollar sign marks the end of the URL. Disallow: /*.pdf$ blocks URLs that end in .pdf. Disallow: /*?sessionid= blocks any URL carrying that parameter. Because matching is case-sensitive, /Photos/ and /photos/ are different paths.
Group selection and rule precedence
Two separate decisions happen when a crawler reads your file: which group applies to it, and which rule inside that group wins for a given URL.
Which group. Only one group is valid for a crawler. Google picks the group whose user-agent line matches most specifically and ignores the rest. The order of groups in the file does not matter. If a file has several groups for the same user agent, Google merges their rules. Specific groups are not merged with the asterisk group. That means a Googlebot group replaces the asterisk group for Googlebot rather than adding to it, so any rules you want Googlebot to keep must be repeated in its own group.
Which rule. Within the group, the most specific rule wins, measured by the length of the path. If an Allow and a Disallow match with equal specificity, Google uses the least restrictive one, which means Allow wins. The table below sets out the points that decide most real-world conflicts.
| Element | How Google treats it |
|---|---|
| Disallow: /folder/ with Allow: /folder/page.html | The longer Allow path is more specific, so the page may be crawled |
| Allow: /folder and Disallow: /folder (equal length) | Conflict resolved to the least restrictive rule, so crawling is allowed |
| Crawl-delay | Not supported by Google, so the line has no effect |
| Robots.txt larger than 500 KiB | Content past the limit is ignored |
| Googlebot group plus an asterisk group | Googlebot follows only its own group; the two are not combined |
AI crawler user agents
Robots.txt is also the tool site owners use to say which AI products may fetch their content. The mechanic is the same as for any crawler: you name a user-agent token and give it rules. What differs is that each company runs several crawlers for different purposes, and the tokens change over time, so confirm the current names on each vendor's own documentation before relying on this list. These are the commonly published ones at the time of writing:
- OpenAI: GPTBot (collection for model training), OAI-SearchBot (search features), ChatGPT-User (fetches triggered by a user action).
- Anthropic: ClaudeBot (training data), Claude-SearchBot (search quality), Claude-User (user-initiated fetches).
- Perplexity: PerplexityBot (its search index) and Perplexity-User (user-triggered).
- Google: Google-Extended is a standalone robots.txt token that controls whether content Google crawls from your site may be used to train future Gemini models and for grounding in Gemini Apps and Vertex AI. It has no separate user-agent string. Google states it does not affect inclusion in Google Search and is not a ranking signal.
- Common Crawl: CCBot, whose public dataset is used by many other organizations.
The decision is a trade-off, not a default. Blocking a training crawler keeps your pages out of that vendor's training collection. Blocking a search or retrieval crawler can mean the assistant cannot cite your pages as a source. Vendors also treat user-triggered fetchers differently, and some state that robots.txt rules may not apply to them, so read the vendor page rather than assuming. Because compliance is voluntary, server logs are the only way to confirm what actually fetched your pages.
The group rules above apply here too. If you write a group for GPTBot, it replaces the asterisk group for that crawler. A group that reads User-agent: GPTBot followed by Disallow: / blocks the whole site for that crawler only.
How to test a robots.txt file
- Fetch it like a crawler would. Request the file with a command-line tool or browser and check three things: the status is 200, the content type is plain text, and the content is the rules you intended rather than an HTML error page. Check every host you run, including the www and non-www versions.
- Use the robots.txt report in Search Console. It shows the robots.txt files Google found for the top 20 hosts of your property, when each was last fetched, and any fetch problems or parse warnings. It is available only for domain-level properties (a Domain property, or a URL-prefix property without a path). Google generally recrawls the file often, so a recrawl request from the report is meant for emergencies, such as a fixed fetch error or a critical rules change.
- Inspect an important URL. The URL Inspection tool has a Crawl allowed? field that shows whether a robots.txt rule blocked Google from crawling the URL. Run it on your top landing pages, your newest articles and a few URLs you intentionally block. All three groups should behave as expected.
- Watch the page indexing report. It lists URLs under statuses such as Blocked by robots.txt and Indexed, though blocked by robots.txt. The second status is the crawl-versus-index trap in action.
- Re-test after every deployment. A robots.txt file is part of the site build, so a release can overwrite it silently.
Common robots.txt mistakes on Thai sites
These are mechanics rather than anecdotes, and each one is cheap to check.
Staging rules shipped to production
Development sites are often blocked with User-agent: * and Disallow: /. If that file travels to the live site during a launch or migration, Google stops crawling everything. Compare the live file to the intended one after every launch.
Blocking CSS and JavaScript
Google renders pages, and its documentation warns that blocking resource files can stop it from understanding pages that depend on them. Block only resources you are sure are unimportant to the rendered page.
Mixing Disallow and noindex
As covered above, a disallowed page cannot show its noindex tag. Choose the tool that matches the goal: Disallow to manage crawling, noindex to keep a page out of results.
Prefix matches that catch more than intended
Because rules match from the start of the path, a short rule can capture unrelated URLs. On a bilingual site where Thai pages live under a /th/ folder, a careless rule such as Disallow: /t would block that folder and every other path beginning with t. Write rules with enough characters to be unambiguous, and test a sample of live URLs against them.
Thai characters in URLs
Thai slugs are common on local sites. Google treats raw UTF-8 characters in a rule and their percent-encoded forms as identical, so either style works. The mistake to avoid is mixing case-sensitive Latin parts of a path, such as /Blog/ versus /blog/, or assuming a rule written for one folder name covers a similar one.
Wrong host, wrong protocol
If your CDN, a staging subdomain or an asset host serves its own robots.txt, the main file does not apply to it. Check each hostname that serves content you care about.
A server that returns an error for the file
A firewall or security plugin that returns 403 or 5xx to crawlers on /robots.txt changes how Google treats the whole site, as the status-code section above explains. Request the file from a few different networks and check the server log for what Googlebot received.
Relative sitemap lines
The Sitemap field needs a full URL with protocol and host. A path alone is invalid.
What this means for Thai businesses
Most Thai business sites run on WordPress, Shopify or another platform that generates a default robots.txt. That default is a reasonable starting point, but platforms add URLs the default does not cover: filtered product listings, internal search pages, tag archives and tracking-parameter duplicates. Whether those deserve a Disallow rule is a judgement about your site size and server capacity, not a universal rule. Small sites with a few hundred URLs rarely need crawl management at all, and an overly aggressive file does more harm than good.
The decision with a longer shelf life is the AI crawler policy. A business that wants to be cited in AI answers has a reason to allow search and retrieval crawlers, even if it declines training crawlers. A business with licensed or paywalled content may decide the opposite. Either way, write the choice down, test it, and review the vendor token list every few months, because the names and purposes change. For a full technical review of crawling, indexing and the robots rules on your own domain, see our SEO audit service, and for the wider search programme, the SEO in Thailand service.
Robots.txt FAQ
Does robots.txt stop a page from appearing in Google?
No, it only stops Google from crawling the page, and the URL can still be indexed if other sites link to it. Use a noindex rule, password protection or removal to keep a page out of search results.
Where does the robots.txt file go?
It goes at the top level of each host, such as https://www.example.com/robots.txt. A file inside a subfolder is not read, and subdomains, other protocols and other ports each need their own file.
Which rule wins when Allow and Disallow conflict?
The rule with the longer, more specific path wins for Google. If two conflicting rules are equally specific, Google uses the least restrictive one, so the Allow rule applies.
How do I block AI crawlers in robots.txt?
Add a group for each crawler's user-agent token, such as GPTBot or ClaudeBot, with Disallow: / to block the whole site for that crawler. Confirm the current token names on each vendor's documentation, and remember that compliance is voluntary.
How long does a robots.txt change take to apply?
Google generally caches the file for up to 24 hours, so changes can take about a day to apply. For a fixed error or a critical change, you can request a recrawl of the file in the Search Console robots.txt report.
If you are unsure whether your own file is helping or hurting, a technical review of the live file, your sitemap and your indexing reports will answer it with evidence. Our team can run that review as part of an SEO audit or fold it into a wider SEO services engagement.







