1. Home
  2. Developer
  3. robots.txt Generator

robots.txt Generator

Build robots.txt rules for search and AI crawlers, and test paths against them.

Runs in your browser · No upload Works offline

All search engines and bots (User-agent: *)

Start with /. * matches anything; a trailing $ marks the end of the path. e.g. /admin/, /*?sort=, /*.pdf$
Opens part of a blocked path, e.g. block /admin/ but allow /admin/help
s
Google ignores this; only some bots such as Bing and Yandex honor it.

Block AI crawlers

AI companies use different bot names for different purposes. Blocking only training bots still lets AI search and answers cite you; blocking search bots too may remove you from AI search results.

AI training — collecting data to train models
AI search — indexing pages as sources for AI search and answers
User requests — fetching a page when a user gives the AI a link

Google-Extended and Applebot-Extended aren't separate crawlers; they control whether content collected by Google's and Apple's search crawlers may be used for AI training. Blocking them doesn't affect search visibility.


Rules for specific bots


Use full URLs starting with https://.
robots.txt
User-agent: * Allow: /admin/help Disallow: /admin/ Disallow: /cart Disallow: /checkout Disallow: /search Disallow: /*?sort= User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: meta-externalagent User-agent: Bytespider Disallow: / Sitemap: https://example.com/sitemap.xml

Test a path

robots.txt is a request, not access control. Search engines and major AI companies' bots follow it, but some bots don't. Protect private pages with a login or server configuration.

How to use

  1. List the paths you don’t want crawled under Paths to block, one per line — admin pages, cart, internal search results. “Quick presets” fill in allow-all, block-all or an example.
  2. Pick bots under Block AI crawlers. AI training bots are selected by default.
  3. Use Add bot-specific rules to give one bot different rules.
  4. Add your sitemap URLs, download the file and upload it to your site root (https://yourdomain/robots.txt).
  5. Under Test a path, enter a bot name and path to see whether the current rules allow it, and which rule decided.

Reading the rules (RFC 9309)

Line Meaning
User-agent: * Which bot the rules below apply to (* = every bot not named elsewhere)
Disallow: /admin/ Don’t crawl URLs starting with /admin/
Allow: /admin/help An exception inside a blocked range
Disallow: (empty) Blocks nothing — everything allowed
Disallow: /*?sort= * matches anything — every URL containing ?sort=
Disallow: /*.pdf$ $ marks the end — only URLs ending in .pdf
Sitemap: https://… Where your sitemap lives (full URL)

A bot follows the group that names it exactly; if none does, it follows the * group. When several rules match a URL, the longest path wins, and on a tie, Allow wins. That’s why blocking /admin/ and allowing /admin/help opens just the help page.

Which AI crawlers should you block?

AI companies split their bots by purpose. OpenAI, for example, uses GPTBot for training, OAI-SearchBot for ChatGPT search, and ChatGPT-User when a user shares a link in a conversation.

  • Block training only (the default): you don’t want your content used to train models, but you’re fine being cited as a source in AI search answers.
  • Block search bots too: you don’t want to appear in AI search or answers at all — which can also mean fewer visits from them.
  • Google-Extended and Applebot-Extended aren’t crawlers that visit your site; they decide whether content collected by Google’s and Apple’s search crawlers may be used for AI training. Blocking them doesn’t affect Google Search or Apple search visibility.

Things to watch out for

  • robots.txt is a request, not a lock. Major search engines and AI companies say they honor it, but some bots ignore it, and anyone can read the file. Listing secret paths actually advertises them — protect private pages with a login.
  • Disallow stops crawling, not indexing: a blocked URL that other sites link to can still appear in results without a description. To keep a page out of search, use a noindex meta tag and leave the page crawlable so bots can see it.
  • Google ignores Crawl-delay and stops reading after 500 KiB.

FAQ

Where does the file go?

At the root of your domain, named robots.txt. https://example.com/robots.txt works; https://example.com/blog/robots.txt is ignored. Each subdomain (shop.example.com) needs its own file.

Can I generate meta tags too?

Create title, description and social sharing tags with the Meta Tag Generator, and check your sitemap XML with the XML Formatter.

Coming soon

Can't find the tool you need?

Tell us what you want to do. We build the most requested tools first.