What the check looks at
For the home page of the site you enter, the check reads robots.txt the way the crawlers do and asks for the page under each crawler's name. Then it looks for five things we have seen hide a site from AI answers in practice.
- Cloudflare's email obfuscation, which leaves crawlers reading "[email protected]" where your address should be.
- Words that only appear after JavaScript runs. The OpenAI, Anthropic and Perplexity crawlers read the HTML without running its scripts.
- A noindex instruction on the home page.
- A sitemap, listed in robots.txt or at /sitemap.xml.
- An llms.txt file, a plain-text summary some AI tools read.
Why a 200 is a good sign and nothing more
The requests come from our server with each crawler's name. Some firewalls also check where a request comes from, so the real crawler can be treated differently. A 403 is worth chasing: when we found one on our own site, it matched a setting that turned the real crawlers away too.
If the site refuses even the browser-named request, the site is turning away the checker, and the crawler results cannot tell you much. robots.txt is still read and judged.
What we keep
We keep each result for ten minutes, so a repeated check does not load the same site again, and a count of checks per visitor for one hour, stored against a one-way hash of the IP address, to stop abuse. Nothing else is kept.
The story behind it: two ways our own sites were hidden from AI.
Nicasia Digital Solutions