Crawlers visit your site all day: search engines, AI training bots, AI assistants and SEO tools. Your robots.txt file tells them what to fetch and what to skip.
Keep bots out of cart, search, admin and duplicate URLs so crawl time goes to the pages you want found.
Opt out of AI training and stay in AI search, or block every AI bot. Each of the 21 bots has its own switch.
A Sitemap line lets search engines and AI crawlers discover your URLs without a separate submission.
yourdomain.com
Your rules decide what crawlers see
What robots.txt cannot do. It is a request, not a lock: well-behaved crawlers follow it and some bots ignore it. It does not hide a page from search results, because a disallowed URL can still appear if other sites link to it. Use a noindex tag or a login for that. And the file is public, so never list secrets in it.
Four quick steps, no coding.
https://yourdomain.com/robots.txt Correct: top level of your domainhttps://yourdomain.com/blog/robots.txt Ignored: subfolders do not counthttps://blog.yourdomain.com/robots.txt Each subdomain needs its own fileSave it as plain UTF-8 text, upload it with your hosting File Manager or FTP, then run the AI Crawler Access Checker to confirm it works.
A robots.txt file is a list of short rules grouped by crawler. Here is what each line does.
User-agent: *Disallow: /admin/Allow: /admin/help.htmlUser-agent: GPTBotDisallow: /Sitemap: https://www.yourdomain.com/sitemap.xmlDisallow: /private also blocks /private-beta/. Add a trailing slash to limit a rule to the folder.
/Admin/ and /admin/ are different paths. Match the exact casing of your URLs.
Google and Bing support * and $ in paths, but they are not part of the core standard.
AI companies run separate crawlers for separate jobs, and each can be controlled on its own. Choose by what you want: stay visible in AI answers, keep content out of training, or both.
Collect content that may be used to train AI models.
If you disallow them: your content is excluded from future training. It does not remove you from AI search.
Build the index behind AI search answers and citations.
If you disallow them: you can disappear from that product's search answers. Most sites should leave these allowed.
Fetch a page only when a person asks an AI assistant about it.
Watch out: OpenAI notes robots.txt rules may not apply to ChatGPT-User, because a person starts each request.
The most common goal: stay citable but opt out of training. Click the Disallow AI training bots preset in the generator above and leave search and assistant bots allowed. Google-Extended only controls Gemini training and does not affect Google Search.
Already uploaded your file? See which AI crawlers it allows or blocks.
Check my AI crawler accessAI companies run separate crawlers for separate jobs, and each can be controlled on its own. Choose by what you want: stay visible in AI answers, keep content out of training, or both.
Collect content that may be used to train AI models.
Build the index behind AI search answers and citations.
Fetch a page only when a person asks an AI assistant about it.
Already uploaded your file? See which AI crawlers it allows or blocks.
Check my AI crawler accessPick a setup, copy or download it, and replace yourdomain.com. Need more control? Build your own in the generator above.
Keeps the admin area out of crawlers' way while allowing the file many themes and plugins rely on.
User-agent: * Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://www.yourdomain.com/wp-sitemap.xml
/wp-admin/admin-ajax.php availableKeeps cart, checkout and account pages out of crawlers' way. The wildcard line is supported by Google and Bing.
User-agent: * Disallow: /cart/ Disallow: /checkout/ Disallow: /my-account/ Disallow: /*?add-to-cart= Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://www.yourdomain.com/wp-sitemap.xml
admin-ajax.php availableOpts out of the main AI training crawlers while leaving AI search and search engines untouched.
User-agent: * Allow: / User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: CCBot Disallow: / Sitemap: https://www.yourdomain.com/sitemap.xml
Asks every crawler to stay away. Use it only on sites you do not want crawled, and remove it before launch.
User-agent: * Disallow: /
A small typo in this file can hide your whole site from search. Check these before you upload.
It is meant for staging. On a live site it blocks every crawler that has no rule of its own, including search engines.
It must load at /robots.txt on the top level of each host. A subfolder copy is ignored, and subdomains need their own.
The file is public, and blocked URLs can still appear in results. Protect private pages with a login, and use noindex to keep pages out of search.
Google needs these files to render your pages properly. Do not disallow folders that hold them.
A rule for /Private will not match /private, and /shop also matches /shop-sale/.
Disallowing OAI-SearchBot or PerplexityBot can remove you from AI answers. To opt out of training, disallow GPTBot and ClaudeBot instead.
From uploading the file to enforcing your rules, get help from a team with 23+ years in web hosting.
Drop robots.txt into your site root with File Manager or FTP. Support answers tickets in under 11 minutes.
Hide cart, checkout and account pages from crawlers on hosting built for WooCommerce.
Our SEO and development team can handle robots.txt, sitemaps and technical fixes for you.
Current host getting in the way? Move to AccuWeb Hosting with free website migration.
Running WordPress? Explore WordPress hosting or VPS hosting for full server control.
Answers about robots.txt syntax, AI bots, sitemaps and uploading your file.
yourdomain.com/robots.txt./blog/robots.txt is ignored, and each subdomain needs its own file. The file must also be UTF-8 text, and Google reads up to 500 KiB of it.Disallow: /, for example User-agent: GPTBot followed by Disallow: /. In the generator, just click the bot to switch it to Disallowed. Use the presets to block all training crawlers or every AI bot in the list with one click.Disallow: / to the User-agent: * group on a live site, because that blocks every crawler that has no rule of its own, including search engines.noindex meta tag, or protect it with a login. Never rely on robots.txt for private content, because the file itself is public./folder/ for a directory or /page.html for a single page. Paths are case-sensitive and match by prefix, so /private also blocks /private-beta/. Adding the trailing slash limits the rule to the folder.https://www.yourdomain.com/sitemap.xml. The generator adds a Sitemap: line, which lets any crawler find your sitemap without a separate submission. You can list more than one, separated by commas.Disallow: /.Crawl-delay directive, while Bing and some other crawlers honour it. To manage Googlebot's crawl rate, use the tools in Google Search Console instead.yourdomain.com/robots.txt in a browser to confirm it loads as plain text. Then run the AI Crawler Access Checker to see which AI crawlers your rules allow or block, and use the robots.txt reports in Google Search Console for Googlebot.Our team is available 24x7x365 to help with your robots.txt and hosting.
See our Cookie Policy