If you ask yourself, "Can ChatGPT read my website?", the answer usually sits in a plain text file hosted at your root domain. AI assistants browse the web to answer questions, pull citations and, through separate crawlers, train models. Running an AI crawler checker on your domain is the quickest way to verify whether your digital storefront is visible to modern answer engines or blocked behind an accidental line of configuration code.
For years, robots.txt was primarily managed to control search engines like Googlebot and Bingbot. Today, an overzealous security plugin or a cautious developer update can silently turn away ChatGPT, Claude, and Perplexity. When an AI crawler cannot access your content, the platform cannot evaluate your product pages, read your service breakdowns, or cite your articles as answers to prospective buyers. Understanding the rules governing these bots is essential for preserving your organic presence in modern search environments, as explored in our guide on how AI search changes traditional SEO.
Understanding the three types of AI bots
Not every bot deployed by an AI company serves the same purpose. Lumping every machine visitor into one category often leads business owners to block everything out of caution, cutting off valuable referral traffic in the process. AI agents generally fall into three functional groups:
- Search and answer crawlers: These automated agents index the web specifically to provide real-time information, links, and source citations inside AI chat interfaces. Blocking these bots removes your business from conversational search results.
- Model training crawlers: These bots gather large datasets of public text across the web to train future foundation models. Disallowing these bots informs the provider that your written material should not be used to build future base models.
- User-initiated fetchers: When a person inside a chat interface pastes your link and asks for a summary, or runs a workflow inside an assistant, a user-directed fetcher visits your domain. Because these requests represent explicit human actions, some platforms do not subject them to standard robots.txt restrictions.
Separating search indexing from model training allows businesses to protect their proprietary writing from generative training while remaining fully discoverable to buyers who search within AI assistants.
AI crawlers compared: purpose and blocking impact
Each platform maintains distinct user-agent tokens for crawling. The table below outlines the primary agents used by OpenAI, Anthropic, Perplexity, and Google.
| Company | Bot user-agent | Primary purpose | Effect of blocking |
|---|---|---|---|
| OpenAI | OAI-SearchBot | Surfaces websites in ChatGPT search features | Your site will not appear in ChatGPT search answers |
| OpenAI | GPTBot | Gathers web content to train foundation models | Content is flagged not to be used for model training |
| OpenAI | ChatGPT-User | Executes user actions in ChatGPT and Custom GPTs | May not follow robots.txt rules because requests are user-initiated |
| Anthropic | Claude-SearchBot | Navigates the web to improve search result quality | Restricts Claude from discovering content for search queries |
| Anthropic | ClaudeBot | Collects web content that could contribute to training | Signals an opt-out from Anthropic model training runs |
| Anthropic | Claude-User | Accesses pages when users ask direct questions | Hinders live browsing actions requested by Claude users |
| Perplexity | PerplexityBot | Discovers and links pages in Perplexity search | Your site is omitted from citations in Perplexity answers |
| Perplexity | Perplexity-User | Fetches content on behalf of specific user queries | Generally ignores robots.txt as an active human query |
| Google-Extended | Controls training for Gemini models and Vertex AI grounding | Content is excluded from model training without altering Google Search rank |
According to technical documentation from OpenAI, a webmaster can easily allow OAI-SearchBot to appear in conversational search results while disallowing GPTBot to restrict training. Similarly, Anthropic notes that its bots respect standard robots.txt signals, allowing site owners to opt out directly via configuration files.
For Perplexity, technical guides state that PerplexityBot surfaces web links and does not train base models, meaning blocking it removes your URLs from answer engine citations. Meanwhile, Google specifies that disallowing Google-Extended has no negative impact on inclusion or ranking within traditional Google Search.
How to test your site with an AI crawler checker
Checking your robots.txt manually requires parsing user-agent blocks and matching wildcard directives line by line. Using an automated AI crawler checker reads your site's robots.txt file under the standard RFC 9309 rules that commercial crawlers obey.
The check inspects 19 separate AI user-agents across two primary groups: 11 search and answer crawlers (such as OAI-SearchBot, Claude-SearchBot, PerplexityBot, and Googlebot) and 8 training bots (including GPTBot, ClaudeBot, Google-Extended, and CCBot). The tool pinpoints the exact line permitting or preventing each bot from reading your pages, and confirms whether your robots.txt points to your sitemap.
Knowing your accessibility status is the first step toward building a predictable workflow for appearing in ChatGPT, Claude, and Perplexity search answers. Note that passing crawler checks is not an absolute guarantee of citation. It simply confirms that technical barriers are not turning bots away at your server door.
How to configure your robots.txt for AI search and training
If you want your content cited when users research products in AI tools, but prefer not to donate your editorial assets to future model training runs, you can create a targeted robots.txt file.
Here is a practical configuration example:
# Allow search engines and AI search assistants
User-agent: Googlebot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# Block AI training bots from scraping content
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
# Sitemap location
Sitemap: https://example.com/sitemap.xml
Deploying this configuration keeps your pages accessible to real-time assistants while placing explicit boundaries around foundation training. Note that robots.txt is a request, not a lock. Anthropic says its bots honour robots.txt, and OpenAI and Perplexity describe it as how you control their crawlers, but changes are not instant: both say it can take up to about 24 hours for their systems to reflect an update.
Why AI assistants still miss your pages when robots.txt is clear
Sometimes an operator permits every crawler in robots.txt, yet AI platforms still fail to access the site. This discrepancy usually stems from edge security platforms, hosting firewalls, or rendering issues:
- CDN and firewall bot protection: Firewalls and CDNs can block crawlers outright, whatever robots.txt says. Cloudflare, for example, has a setting that blocks AI crawlers on every site that turns it on. If the checker says you are open but assistants never mention you, look there next.
- IP-level blocking: Anthropic warns that attempting to restrict crawler behaviour by blocking IP subnets at your firewall can backfire. If your server bans an IP range entirely, the crawler cannot read your robots.txt file to discover your actual permissions.
- JavaScript rendering limits: Some crawlers may not run client-side scripts, so check that your key content is in the HTML the server returns.
Review your CDN security logs for 403 Forbidden or 429 Too Many Requests responses served to user-agents like OAI-SearchBot or ClaudeBot to ensure legitimate search crawlers are not being snagged by automated security shields.
A sensible default strategy for growing businesses
For modern businesses, total isolation is rarely practical. If your revenue depends on organic customer discovery, blocking every automated agent will gradually erode your visibility as search behaviour migrates into conversational engines.
A balanced posture for commercial websites:
- Permit real-time search bots: Explicitly allow OAI-SearchBot, Claude-SearchBot, and PerplexityBot alongside standard search crawlers like Googlebot.
- Decide on training scrapers: If you publish proprietary research or creative work, feel free to disallow GPTBot and ClaudeBot. Blocking training crawlers does not stop the search and answer crawlers reading your pages.
- Verify complete technical health: Checking robots.txt is only the beginning. Validate your XML links with a sitemap checker, review canonical and meta robots tags using an indexability checker, and structure your site summaries using an llms.txt checker. You can browse the complete suite of free web utility tools to audit your technical foundation without an account.
For businesses that want ongoing search insight, Ergora is an AI business suite with seat plans from $99 a month and specialist packs from $59 a month, with 50% off the first month and a 30-day money-back guarantee. Once you connect Google Search Console to Ergora, the SEO specialist can read it: the queries you rank for, pages losing clicks, pages with high impressions but few clicks, and whether a page is in Google's index.
Frequently asked questions
Can ChatGPT read my website?
Yes, ChatGPT can read your website as long as your robots.txt file does not disallow its search bots and your web server does not block its access. OpenAI says OAI-SearchBot is what surfaces websites in ChatGPT's search features, while ChatGPT-User fetches pages when a person asks ChatGPT to look at one. If your site blocks this crawler or relies on complex scripts that require browser execution, ChatGPT will be unable to retrieve your information.
Should I block AI crawlers?
Whether to block AI crawlers depends entirely on your business model and content goals. If you want your products, services, and articles recommended by conversational search engines, you should allow search bots like OAI-SearchBot and PerplexityBot. If you want to protect proprietary writing from being absorbed into future foundation training runs, you can block training bots like GPTBot, ClaudeBot, and Google-Extended while keeping search access intact.
What is the difference between GPTBot and OAI-SearchBot?
OpenAI operates separate crawlers for training and search. GPTBot collects web data specifically to train generative foundation models, and disallowing it signals that your content should be excluded from future training datasets. In contrast, OAI-SearchBot indexes pages to display source links and citations within ChatGPT search, meaning blocking it removes your domain from conversational search answers.
Does blocking Google-Extended affect my Google rankings?
No, blocking Google-Extended does not hurt your positions or visibility in regular Google Search. Google-Extended is a dedicated token that dictates whether your content can be used to train Gemini models and ground responses in Vertex AI. Google says it does not impact a site's inclusion in Google Search nor is it used as a ranking signal.
How do I block ClaudeBot in robots.txt?
To stop Anthropic's training crawler from accessing your site, add a dedicated user-agent directive to your robots.txt file. Include User-agent: ClaudeBot followed immediately by Disallow: / on the next line. Anthropic says its bots honour robots.txt directives.
Why can AI tools not see my site when robots.txt allows everything?
Even if your robots.txt file permits all bots, network-level security rules can still block access. Firewalls and CDNs can block crawlers outright, and Cloudflare, for example, has a setting that blocks AI crawlers on every site that turns it on. Some crawlers may also not run client-side JavaScript, so check that your key content is in the HTML the server returns.