Skip to content
wellread.Run free audit
Wellread/Guides

How to check that AI crawlers can read your website

Last updated 8 October 2026

An AI assistant can only cite a page it can read. This guide is the checklist we use, in the order that finds problems fastest. We wrote it after finding that our own website, built as a React app, served an empty page to any reader that doesn’t run JavaScript.

The short version

  1. Fetch a page without a browser and look for a sentence you can see on it. If it isn’t there, some AI crawlers get an empty page.
  2. Read your robots.txt for the AI search crawlers, path by path.
  3. Check that your CDN or firewall isn’t turning those crawlers away, whatever robots.txt says.
  4. Check each page for noindex and nosnippet, and that a made-up address answers 404.
  5. Check each page names its own address as canonical.
  6. Ask the search engines themselves whether they have the page. Reading is not the same as indexing.

What happened to us

Until 7 October 2026, every public address on trywellread.com returned the same 3,378 bytes of HTML: a title, a description and an empty container that JavaScript filled in the browser. A reader that ran no JavaScript got 0 words. An address that didn’t exist answered 200 with the same shell, and every page told search engines that its canonical address was the home page.

We changed the build so each public page is written out in full as HTML. The home page then arrived with about 1,000 words in it, and each page has its own title, description and canonical address. An address with no page answers 404.

Two things we learned on the way are worth passing on. First, Google had indexed all four pages anyway, because Google’s crawler runs JavaScript: “readable by Google” and “readable by every AI crawler” are different tests. Second, fixing it changed nothing in what the assistants said the day after. Our own audit has the numbers.

1. Fetch the page without a browser

Pick a sentence you can see on the page and look for it in what the server sends:

curl -s https://yourcompany.com/pricing | grep -c "a sentence from that page"

A count of 0 means the sentence wasn’t in what the server sent. Rule out the simple causes first: a redirect (add -L), an error status, or wording that differs from what you typed. If the page loads in a browser and the sentence still isn’t in the response, it is being added by JavaScript after the page loads. A study by Vercel and MERJ of crawler requests, published in December 2024, found that “none of the major AI crawlers currently render JavaScript”, naming those of OpenAI, Anthropic and Perplexity among others; Google’s crawler, which Gemini relies on, does (The rise of the AI crawler). That was measured then; nobody outside those companies can say it for every crawler today. What is safe to say: a page whose words are in the HTML can be read by a crawler either way.

Do this for the pages that matter, not only the home page: pricing, the product page, comparisons, documentation.

2. Read robots.txt for the right crawlers

curl -s https://yourcompany.com/robots.txt

Each AI company runs separate crawlers for separate jobs. Blocking one is not blocking the others:

CompanySearch (answers)Fetching for a userTraining
OpenAIOAI-SearchBotChatGPT-UserGPTBot
AnthropicClaude-SearchBotClaude-UserClaudeBot
PerplexityPerplexityBotPerplexity-UserNone documented

OpenAI says that sites opted out of OAI-SearchBot “will not be shown in ChatGPT search answers”, and that a robots.txt change can take about 24 hours to take effect (OpenAI’s crawlers). Blocking GPTBot or ClaudeBot is a choice about training. It does not keep you out of answers, and allowing them is not a way into answers. A fetch made for a user is different again: Perplexity says Perplexity-User “generally ignores robots.txt rules”, because a person asked for that page. See also Anthropic’s crawlers and Perplexity’s crawlers.

Google-Extended is different again: it is a name in robots.txt, not a crawler. It controls whether Google may use your pages for Gemini training and grounding, and it has no effect on Google Search (Google’s crawlers).

Three rules of the robots.txt standard (RFC 9309) trip people up:

  • The longest matching rule wins. Disallow: /app/ blocks the /app/ folder. It is not Disallow: /, which blocks everything. A checking tool once told us our site blocked all bots; it had read the first as the second.
  • A group that names a crawler replaces the general group for that crawler. If you write rules for OAI-SearchBot, the rules under * no longer apply to it.
  • A robots.txt that answers with a server error closes the whole site. Crawlers that follow the standard must then assume everything is disallowed. A missing file (404) means everything is allowed.

3. Check the firewall, not only the file

robots.txt is a request, and your CDN or firewall is the door. Bot protection can refuse a crawler that robots.txt allows. A quick test is to ask as the crawler:

curl -s -o /dev/null -w "%{http_code}\n" -A "OAI-SearchBot" https://yourcompany.com/

A 403, a 429 or a challenge page here is worth chasing. A 200 is weaker evidence than it looks: your request came from your computer with the crawler’s name on it, and the real crawler comes from its company’s own addresses, which a firewall can treat differently. The real test is in your CDN’s logs: look for the crawler’s requests and the status they got. OpenAI, Anthropic and Perplexity each publish their crawlers’ addresses for this.

4. Check what each page tells search engines

  • noindex and nosnippet. Google says that to appear as a supporting link in AI Overviews or AI Mode, a page “must be indexed and eligible to be shown in Google Search with a snippet” (AI features and your website). Look in the page’s HTML for a robots meta tag, and in its response for an X-Robots-Tag header.
  • Missing pages. An address that doesn’t exist should answer 404, not 200 with your home page:
    curl -s -o /dev/null -w "%{http_code}\n" https://yourcompany.com/no-such-page
  • Canonical address. Each page should name its own address in its canonical tag. A single-page app often sends the home page’s tag on every page, as ours did.

5. Ask the search engines whether they have it

Everything above tests whether a page can be read. Whether it has been indexed is a separate fact that only the search engine knows. In Google Search Console, URL inspection says whether a page is on Google, when it was last crawled and which address Google chose as canonical. Bing Webmaster Tools does the same for Bing. Submit your sitemap in both. A new site can be perfectly readable and still not be in an index yet.

How to fix a page that arrives empty

  • Write the pages out when the site is built. This suits pages that are the same for everyone: home, pricing, guides. It is what we did, with React’s own server renderer and the Vite build we already had. No new framework was needed.
  • Render on the server for each request, when a public page’s content changes per request.
  • Keep the app as it is. Only the public pages need this. A signed-in product can stay a browser app.

Don’t build a second version of a page for crawlers that says something different from what people see. Serve everyone the same HTML and let the browser take it over.

What these checks don’t prove

  • That the real crawlers get through. Your tests come from your own address.
  • That a page is indexed. Only the search engine’s own tools say.
  • That an assistant will cite or recommend you. Being readable is the entry ticket, not the result.

Wellread runs these checks on up to 60 pages as part of every audit, then asks ChatGPT and Gemini your buyers’ questions. What an audit is.

Written by the Wellread team. If something here is wrong or out of date, tell us at support@trywellread.com and we will fix it.

wellread.

Find out whether AI recommends you, and what to change.

How it worksThe auditSample reportPricingHow we measureGuidesAboutFAQPrivacyTermssupport@trywellread.com© 2026 Wellread