AccuWeb Hosting

VPS Hosting

VIRTUAL PRIVATE SERVERS
Full root access and guaranteed resources.
APPS ON A VPS
Deploy popular applications on your own server.
All applications

More Services

AI Website & Store Builder
Build a website or online store with AI.
Start
WEB & WORDPRESS HOSTING
For websites, blogs and online stores.
RESELLER & AGENCY
DOMAINS, SECURITY & EMAIL
Domain names, protection and professional email.
  1. Home
  2. Robots.txt Generator

Robots.txt Generator Tool

Create a valid, SEO-friendly robots.txt file in seconds without writing manual syntax.

Search Engines

4
Googlebot
Bingbot
YandexBot
DuckDuckBot

AI Bots

4
GPTBot
ClaudeBot
Google-Extended
CCBot

SEO Tools

3
AhrefsBot
SemrushBot
MJ12bot

Social Media

3
FacebookExternalHit
Twitterbot
LinkedInBot

Custom Bots

0

Restricted Directories

0

Add paths you want to block for all bots (e.g. /admin/, /private/)

Copied successfully!
Why it matters

Why every website needs a robots.txt file

Crawlers visit your site all day: search engines, AI training bots, AI assistants and SEO tools. Your robots.txt file tells them what to fetch and what to skip.

  • Focus crawlers on pages that matter

    Keep bots out of cart, search, admin and duplicate URLs so crawl time goes to the pages you want found.

  • Decide what AI companies can use

    Opt out of AI training and stay in AI search, or block every AI bot. Each of the 21 bots has its own switch.

  • Help every crawler find your sitemap

    A Sitemap line lets search engines and AI crawlers discover your URLs without a separate submission.

What robots.txt cannot do. It is a request, not a lock: well-behaved crawlers follow it and some bots ignore it. It does not hide a page from search results, because a disallowed URL can still appear if other sites link to it. Use a noindex tag or a login for that. And the file is public, so never list secrets in it.

How it works

How to create and upload your robots.txt file

Four quick steps, no coding.

Where your robots.txt file must live

https://yourdomain.com/robots.txt Correct: top level of your domain
https://yourdomain.com/blog/robots.txt Ignored: subfolders do not count
https://blog.yourdomain.com/robots.txt Each subdomain needs its own file

Save it as plain UTF-8 text, upload it with your hosting File Manager or FTP, then run the AI Crawler Access Checker to confirm it works.

Syntax guide

Robots.txt syntax explained: User-agent, Disallow, Allow and Sitemap

A robots.txt file is a list of short rules grouped by crawler. Here is what each line does.

Paths match by prefix

Disallow: /private also blocks /private-beta/. Add a trailing slash to limit a rule to the folder.

Paths are case-sensitive

/Admin/ and /admin/ are different paths. Match the exact casing of your URLs.

Wildcards are widely supported

Google and Bing support * and $ in paths, but they are not part of the core standard.

AI bots

Which AI bots should you disallow in robots.txt?

AI companies run separate crawlers for separate jobs, and each can be controlled on its own. Choose by what you want: stay visible in AI answers, keep content out of training, or both.

Your choice

Training crawlers

Collect content that may be used to train AI models.

GPTBot ClaudeBot Google-Extended CCBot

If you disallow them: your content is excluded from future training. It does not remove you from AI search.

Usually keep allowed

Search crawlers

Build the index behind AI search answers and citations.

OAI-SearchBot Claude-SearchBot PerplexityBot

If you disallow them: you can disappear from that product's search answers. Most sites should leave these allowed.

Rules may not apply

Assistant crawlers

Fetch a page only when a person asks an AI assistant about it.

ChatGPT-User Claude-User Perplexity-User

Watch out: OpenAI notes robots.txt rules may not apply to ChatGPT-User, because a person starts each request.

The most common goal: stay citable but opt out of training. Click the Disallow AI training bots preset in the generator above and leave search and assistant bots allowed. Google-Extended only controls Gemini training and does not affect Google Search.

AI training crawlers (8)

  • GPTBot (OpenAI)
  • ClaudeBot (Anthropic)
  • Google-Extended (Google)
  • Applebot-Extended (Apple)
  • meta-externalagent (Meta)
  • CCBot (Common Crawl)
  • Bytespider (ByteDance)
  • cohere-ai (Cohere)

AI search crawlers (4)

  • OAI-SearchBot (OpenAI)
  • Claude-SearchBot (Anthropic)
  • PerplexityBot (Perplexity)
  • DuckAssistBot (DuckDuckGo)

AI assistant crawlers (5)

  • ChatGPT-User (OpenAI)
  • Claude-User (Anthropic)
  • Perplexity-User (Perplexity)
  • MistralAI-User (Mistral AI)
  • Meta-ExternalFetcher (Meta)

Other AI crawlers and agents (4)

  • Amazonbot (Amazon)
  • DeepSeekBot (DeepSeek)
  • GrokBot (xAI)
  • AI2Bot (Allen Institute for AI)

Already uploaded your file? See which AI crawlers it allows or blocks.

Check my AI crawler access
AI bots

Which AI bots should you disallow in robots.txt?

AI companies run separate crawlers for separate jobs, and each can be controlled on its own. Choose by what you want: stay visible in AI answers, keep content out of training, or both.

Your choice

Training crawlers

Collect content that may be used to train AI models.

GPTBot ClaudeBot Google-Extended CCBot
If you disallow them: your content is excluded from future training. It does not remove you from AI search.
Usually keep allowed

Search crawlers

Build the index behind AI search answers and citations.

OAI-SearchBot Claude-SearchBot PerplexityBot
If you disallow them: you can disappear from that product's search answers. Most sites should leave these allowed.
Rules may not apply

Assistant crawlers

Fetch a page only when a person asks an AI assistant about it.

ChatGPT-User Claude-User Perplexity-User
Watch out: OpenAI notes robots.txt rules may not apply to ChatGPT-User, because a person starts each request.
The most common goal: stay citable but opt out of training. Click the Disallow AI training bots preset in the generator above and leave search and assistant bots allowed. Google-Extended only controls Gemini training and does not affect Google Search.

AI training crawlers (8)

  • GPTBot (OpenAI)
  • ClaudeBot (Anthropic)
  • Google-Extended (Google)
  • Applebot-Extended (Apple)
  • meta-externalagent (Meta)
  • CCBot (Common Crawl)
  • Bytespider (ByteDance)
  • cohere-ai (Cohere)

AI search crawlers (4)

  • OAI-SearchBot (OpenAI)
  • Claude-SearchBot (Anthropic)
  • PerplexityBot (Perplexity)
  • DuckAssistBot (DuckDuckGo)

AI assistant crawlers (5)

  • ChatGPT-User (OpenAI)
  • Claude-User (Anthropic)
  • Perplexity-User (Perplexity)
  • MistralAI-User (Mistral AI)
  • Meta-ExternalFetcher (Meta)

Other AI crawlers and agents (4)

  • Amazonbot (Amazon)
  • DeepSeekBot (DeepSeek)
  • GrokBot (xAI)
  • AI2Bot (Allen Institute for AI)

Already uploaded your file? See which AI crawlers it allows or blocks.

Check my AI crawler access
Copy-ready examples

Robots.txt examples you can copy

Pick a setup, copy or download it, and replace yourdomain.com. Need more control? Build your own in the generator above.

WordPress site

Keeps the admin area out of crawlers' way while allowing the file many themes and plugins rely on.

robots.txt
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://www.yourdomain.com/wp-sitemap.xml
  • Blocks /wp-admin/
  • Keeps admin-ajax.php available
  • Points to your WordPress sitemap

WooCommerce store

Keeps cart, checkout and account pages out of crawlers' way. The wildcard line is supported by Google and Bing.

robots.txt
User-agent: *
Disallow: /cart/
Disallow: /checkout/
Disallow: /my-account/
Disallow: /*?add-to-cart=
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://www.yourdomain.com/wp-sitemap.xml
  • Blocks cart, checkout and account pages
  • Blocks add-to-cart URLs
  • Keeps admin-ajax.php available

Stay citable, skip AI training

Opts out of the main AI training crawlers while leaving AI search and search engines untouched.

robots.txt
User-agent: *
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

Sitemap: https://www.yourdomain.com/sitemap.xml
  • Disallows GPTBot, ClaudeBot, Google-Extended and CCBot
  • Search engines and AI search stay allowed

Staging or private site

Asks every crawler to stay away. Use it only on sites you do not want crawled, and remove it before launch.

robots.txt
User-agent: *
Disallow: /
  • Disallows the whole site for all crawlers
  • Not a security measure: add a login
  • Remove it before you go live
Avoid these

Common robots.txt mistakes to avoid

A small typo in this file can hide your whole site from search. Check these before you upload.

Leaving Disallow: / on a live site

It is meant for staging. On a live site it blocks every crawler that has no rule of its own, including search engines.

Putting the file in the wrong place

It must load at /robots.txt on the top level of each host. A subfolder copy is ignored, and subdomains need their own.

Using it to hide private pages

The file is public, and blocked URLs can still appear in results. Protect private pages with a login, and use noindex to keep pages out of search.

Blocking CSS and JavaScript files

Google needs these files to render your pages properly. Do not disallow folders that hold them.

Forgetting paths are case-sensitive prefixes

A rule for /Private will not match /private, and /shop also matches /shop-sale/.

Blocking a search bot when you meant a training bot

Disallowing OAI-SearchBot or PerplexityBot can remove you from AI answers. To opt out of training, disallow GPTBot and ClaudeBot instead.

Go further

Put your robots.txt to work with AccuWeb Hosting

From uploading the file to enforcing your rules, get help from a team with 23+ years in web hosting.

Live on your site today

Drop robots.txt into your site root with File Manager or FTP. Support answers tickets in under 11 minutes.

Get cPanel hosting

Crawl-ready WooCommerce

Hide cart, checkout and account pages from crawlers on hosting built for WooCommerce.

WooCommerce plans

Skip the setup

Our SEO and development team can handle robots.txt, sitemaps and technical fixes for you.

Get expert help

Switch hosts for free

Current host getting in the way? Move to AccuWeb Hosting with free website migration.

Migrate for free

Running WordPress? Explore WordPress hosting or VPS hosting for full server control.

Supporting Over 100K+ Satisfied Businesses

Our customers say Excellent Our customers 4.3 out of 5 based on 217 reviews (TrustPilot)
▶ Hear from our customers
Read all Reviews
FAQ

Robots.txt generator: frequently asked questions

Answers about robots.txt syntax, AI bots, sitemaps and uploading your file.

A robots.txt file is a plain-text file at the root of your website that tells crawlers which URLs they may request. It follows the Robots Exclusion Protocol, which is now standardised as RFC 9309. Search engines, AI crawlers and SEO tools read it before crawling your pages.
Use the generator above. Enter your website and sitemap URL, switch off any AI bots you want to disallow, add custom bots or page and folder paths, and watch the file update live. Then copy the text or download the file and upload it to your site's root folder so it loads at yourdomain.com/robots.txt.
In the top-level directory of your site. Google states that the file must sit at the top level, that its URL is case-sensitive, and that its rules apply only to the host, protocol and port where it is served. A file inside a subfolder such as /blog/robots.txt is ignored, and each subdomain needs its own file. The file must also be UTF-8 text, and Google reads up to 500 KiB of it.
Add a separate group for each bot with Disallow: /, for example User-agent: GPTBot followed by Disallow: /. In the generator, just click the bot to switch it to Disallowed. Use the presets to block all training crawlers or every AI bot in the list with one click.
Blocking GPTBot, ClaudeBot or other AI crawlers does not block Googlebot, which handles Google Search. Google-Extended is a separate control token for Gemini training and is not Googlebot. Never add Disallow: / to the User-agent: * group on a live site, because that blocks every crawler that has no rule of its own, including search engines.
Training crawlers such as GPTBot and ClaudeBot collect content that may be used to train models. Search crawlers such as OAI-SearchBot and PerplexityBot build the index behind AI search answers. Assistant crawlers such as ChatGPT-User and Claude-User fetch a page when a person asks about it. Each can be controlled separately. To stay citable but opt out of training, disallow the training group and leave the other two allowed.
No. robots.txt controls crawling, not indexing. A disallowed URL can still appear in search results if other sites link to it. To keep a page out of search results, leave it crawlable and add a noindex meta tag, or protect it with a login. Never rely on robots.txt for private content, because the file itself is public.
Type the path in the generator and click Disallow path. Use /folder/ for a directory or /page.html for a single page. Paths are case-sensitive and match by prefix, so /private also blocks /private-beta/. Adding the trailing slash limits the rule to the folder.
Enter the full sitemap URL in the Sitemap field, for example https://www.yourdomain.com/sitemap.xml. The generator adds a Sitemap: line, which lets any crawler find your sitemap without a separate submission. You can list more than one, separated by commas.
No. robots.txt is a voluntary standard. Well-behaved crawlers follow it, but some bots ignore it, and OpenAI notes that robots.txt rules may not apply to ChatGPT-User because a person starts those requests. To enforce limits on unwanted bots, use network-level protection such as Blackwall SmartProtect.
Type its user-agent name in the Custom bot field and click Disallow bot. You can find the name in your server access logs or in the bot's documentation. The generator adds a group with that name and Disallow: /.
No. Google ignores the Crawl-delay directive, while Bing and some other crawlers honour it. To manage Googlebot's crawl rate, use the tools in Google Search Console instead.
Open yourdomain.com/robots.txt in a browser to confirm it loads as plain text. Then run the AI Crawler Access Checker to see which AI crawlers your rules allow or block, and use the robots.txt reports in Google Search Console for Googlebot.
Yes, the generator is free. It runs in your browser, so the website, sitemap and rules you enter are used only to build your file on your screen.

Still have questions?

Our team is available 24x7x365 to help with your robots.txt and hosting.

Talk to an expert