Firm Beacon

How to review robots.txt for AI crawlers

Review robots.txt for AI crawlers by listing the user agents that matter, reading their groups and wildcard interactions, testing exact public and private paths, and verifying the live file after deployment. Keep permission separate from visibility.

Updated 7 October 2026. Keep the output and the verification step together.

Steps

  1. Write the policy outcome in plain language: search access, training control or both.
  2. List named user agents from current provider documentation.
  3. Read specific and wildcard groups together.
  4. Test the exact public URL and a private path.
  5. Publish the smallest change and fetch the live file.
  6. Record what remains unproven, including firewall access, indexing and citations.

Start with current documentation

Provider names and purposes are not interchangeable. Keep the source and date beside the policy so a future review can tell what decision the rule was meant to express.

Run a regression check across markets

For a US and European site, test one public URL and one private URL on each host. Compare the named group, wildcard group, longer path and any Allow rule. Record differences instead of collapsing them into one site-wide result. This review is about finding policy drift after a deployment.

Test path matching

A page-level check can reveal that a homepage and a product directory have different outcomes. Check longer paths, Allow rules and a specific crawler group rather than reading only the wildcard block.

Separate infrastructure

If the rule allows a crawler but the request still fails, review authentication, CDN, WAF, DNS and TLS. A robots result cannot diagnose those systems.

Keep the change reversible

Store the previous file, publish through the normal deployment path and rerun the same checks. A narrow, dated change is easier to review than a broad replacement.

Record the standards reference

Use the Google robots.txt documentation for matching rules and the OpenAI bot documentation for current bot names. Reviewed 7 October 2026.

Questions

Can the checker prove a real crawler visit?

No. It evaluates the public robots rules for the entered URL.

Do AI crawlers ignore robots.txt?

Policies and enforcement vary. Treat robots.txt as a requested crawl policy and consult provider documentation for exceptions.

Related resources