How to use
- List the paths you don’t want crawled under Paths to block, one per line — admin pages, cart, internal search results. “Quick presets” fill in allow-all, block-all or an example.
- Pick bots under Block AI crawlers. AI training bots are selected by default.
- Use Add bot-specific rules to give one bot different rules.
- Add your sitemap URLs, download the file and upload it to your site root (
https://yourdomain/robots.txt). - Under Test a path, enter a bot name and path to see whether the current rules allow it, and which rule decided.
Reading the rules (RFC 9309)
| Line | Meaning |
|---|---|
User-agent: * |
Which bot the rules below apply to (* = every bot not named elsewhere) |
Disallow: /admin/ |
Don’t crawl URLs starting with /admin/ |
Allow: /admin/help |
An exception inside a blocked range |
Disallow: (empty) |
Blocks nothing — everything allowed |
Disallow: /*?sort= |
* matches anything — every URL containing ?sort= |
Disallow: /*.pdf$ |
$ marks the end — only URLs ending in .pdf |
Sitemap: https://… |
Where your sitemap lives (full URL) |
A bot follows the group that names it exactly; if none does, it follows the * group. When several rules match a URL, the longest path wins, and on a tie, Allow wins. That’s why blocking /admin/ and allowing /admin/help opens just the help page.
Which AI crawlers should you block?
AI companies split their bots by purpose. OpenAI, for example, uses GPTBot for training, OAI-SearchBot for ChatGPT search, and ChatGPT-User when a user shares a link in a conversation.
- Block training only (the default): you don’t want your content used to train models, but you’re fine being cited as a source in AI search answers.
- Block search bots too: you don’t want to appear in AI search or answers at all — which can also mean fewer visits from them.
Google-ExtendedandApplebot-Extendedaren’t crawlers that visit your site; they decide whether content collected by Google’s and Apple’s search crawlers may be used for AI training. Blocking them doesn’t affect Google Search or Apple search visibility.
Things to watch out for
- robots.txt is a request, not a lock. Major search engines and AI companies say they honor it, but some bots ignore it, and anyone can read the file. Listing secret paths actually advertises them — protect private pages with a login.
- Disallow stops crawling, not indexing: a blocked URL that other sites link to can still appear in results without a description. To keep a page out of search, use a
noindexmeta tag and leave the page crawlable so bots can see it. - Google ignores
Crawl-delayand stops reading after 500 KiB.
FAQ
Where does the file go?
At the root of your domain, named robots.txt. https://example.com/robots.txt works; https://example.com/blog/robots.txt is ignored. Each subdomain (shop.example.com) needs its own file.
Can I generate meta tags too?
Create title, description and social sharing tags with the Meta Tag Generator, and check your sitemap XML with the XML Formatter.