Nicasia Digital Solutions

Field notes · Growth engine

Can ChatGPT and Claude read your website? How to check

Our robots.txt welcomed every crawler, yet Cloudflare sent GPTBot a 403 and hid our email from AI crawlers. What each crawler does, and a one-line test.

On 19 September 2026 we tested one of our own websites the way an AI crawler meets it. Its robots.txt welcomed every crawler. Googlebot, Bingbot, OAI-SearchBot and PerplexityBot got the page. GPTBot, ClaudeBot and CCBot got 25 bytes of plain text from Cloudflare instead: "Your request was blocked."

On 30 September we found a quieter gap on this site, nicasiadigital.com. Every crawler could open the pages, but our email address reached them as "[email protected]".

We had switched on neither setting. Both came with Cloudflare, and neither shows when you open the site in a browser: a browser gets the page, and it runs the script that puts the address back.

Which crawler does what

The companies behind the AI assistants each run more than one crawler, and each crawler feeds something different. Blocking one can cost you a great deal or nothing, depending on which one it is. This is what each company's own documentation says, as of September 2026.

CrawlerCompanyWhat it reads your site for
OAI-SearchBotOpenAIShowing your pages in ChatGPT search answers
GPTBotOpenAITraining future models
ChatGPT-UserOpenAIOpening a page because a user asked ChatGPT to
Claude-SearchBotAnthropicImproving Claude's search results
ClaudeBotAnthropicTraining future models
Claude-UserAnthropicOpening a page because a user asked Claude to
PerplexityBotPerplexityListing and linking sites in Perplexity's answers; Perplexity says it does not train on them
Perplexity-UserPerplexityOpening a page during an answer; it generally ignores robots.txt
Google-ExtendedGoogleA name in robots.txt with no crawler behind it: it decides whether Gemini may train on your pages or use them to ground answers
CCBotCommon CrawlAn open archive of the web that AI companies have used to train models

Two points in that documentation matter most. OpenAI's page says a site that blocks OAI-SearchBot stops appearing in ChatGPT search answers, while blocking GPTBot only keeps its pages out of training. Google's says Google-Extended has no effect on Google Search.

So on our site in September, ChatGPT search could still read us, because OAI-SearchBot got through. What we were losing was training: the next models would know less about the site.

Test it from your own computer

This loop asks for your home page under each crawler's name and prints the status code. It runs in the Mac Terminal and on Linux; on Windows, use Git Bash. Replace the address with yours.

for b in Googlebot bingbot GPTBot \
    OAI-SearchBot ChatGPT-User ClaudeBot \
    Claude-SearchBot Claude-User \
    PerplexityBot CCBot; do
  echo "$b $(curl -s -o /dev/null \
    -w '%{http_code}' -A \
    "Mozilla/5.0 (compatible; $b/1.0)" \
    https://www.example.com/)"
done

On 30 September at 11:05, Malaysia time, nicasiadigital.com answered 200 to all ten names. A 200 means the crawler gets the page. A 403 or a 503 means something in front of your site refuses it, whatever robots.txt says. If you would rather not open a terminal, our free crawler check runs the same test from a web page, along with the checks further down this note.

The test has one limit. It sends each crawler's name from your computer, and some firewalls also check where a request comes from, so the real crawler can be treated differently. Take a 200 as a good sign. Take a 403 seriously: on our site in September it matched a setting that turns away the real crawlers too.

Where Cloudflare keeps the switch

Since 1 July 2025, Cloudflare asks every new domain at sign-up whether to allow AI crawlers, and blocking is where it starts. Our domain moved to Cloudflare on 2 September 2026. Seventeen days later the training crawlers were still being turned away.

In the dashboard, open the domain, then Security, then Settings, and find the Bot traffic group. "Configure AI bot policies" has three rows: Search, Agent and Training. Ours read Allow, Allow, Disallow. Training was the row sending 403 to GPTBot, ClaudeBot and CCBot. We set it to Allow at 15:10 on 19 September, and the next run of the loop showed 200 for all nine names we tested that day.

At the bottom of the same panel, "Enable Bot Preference Sync" lets Cloudflare write its own lines into your robots.txt. We keep it off. The site has a robots.txt of its own, and two sets of rules would sooner or later disagree.

The email address crawlers could not read

Cloudflare's Email Address Obfuscation was on for this domain, and we had not turned it on. It replaces each email address in a page with a link to /cdn-cgi/l/email-protection and the words "[email protected]". A small script then puts the real address back in the visitor's browser, so people never see the swap.

A crawler that reads the HTML and stops there gets "[email protected]". A December 2024 study by Vercel and MERJ found that GPTBot, ClaudeBot and PerplexityBot fetch pages without running their JavaScript, so for them the address is simply gone.

We left the setting on and changed the pages. Cloudflare skips anything between <!--email_off--> and <!--/email_off-->, so our build now wraps every address in those two comments. Turning the feature off for the whole domain works too. Addresses inside <script> tags and in the <head> are never rewritten, which is why the address in our structured data was readable all along.

To check yours, fetch the contact page the way a crawler does and count Cloudflare's markers:

curl -s https://www.example.com/contact \
  -A "Mozilla/5.0 (compatible; GPTBot/1.0)" \
  | grep -c "email-protection"

Anything above 0 means at least one address on that page is hidden from crawlers. On our contact page it now prints 0.

What else a crawler misses

The same study found that Gemini, through Googlebot, and AppleBot do render JavaScript. The OpenAI, Anthropic and Perplexity crawlers it measured did not. Anything your site adds after its scripts run is missing for them: prices loaded by a script, a menu built in the browser, reviews from a widget, sometimes the whole page. To see what they see, fetch the page with curl and search the output for the words that matter to you.

Our checklist after every deploy

  • robots.txt allows the crawlers you want. Under User-agent: * with Allow: /, a crawler that is not named counts as allowed.
  • Every name in the loop above gets 200.
  • On Cloudflare, the AI bot policies say the same thing as robots.txt.
  • Your main words, prices and contact details are in the HTML before any script runs.
  • Your email address prints in the curl check.
  • The sitemap lists every page, with dates that change only when a page's words change.

We run these on nicasiadigital.com after every deploy, from a script that lists anything that fails. If you want the same checks run on your site, a private review covers them.

Sources

More field notes

Start with one form

Tell us where your team's hours go. Your written review arrives within 24 hours.

Request a private review