Firm Beacon

How to block AI training crawlers with robots.txt

To express a preference against a named AI training crawler, review its user-agent group, disallow intended paths and verify the live robots.txt file. This does not automatically block AI search.

Updated 7 October 2026. Written for website owners and teams in the US and Europe.

Direct answer

To express a preference against a named AI training crawler, review its user-agent group, disallow intended paths and verify the live robots.txt file. This does not automatically block AI search.

Name the crawler you mean

Do not assume that every AI system uses one shared user agent. OpenAI documents GPTBot for training and OAI-SearchBot for ChatGPT search. Other providers publish their own names. Keep a record of the documentation that informed your rule.

Decide what training policy covers

For a US product site and a European documentation site, decide whether the preference applies to both public hosts, only selected directories or one market's licensed material. Record the owner and the exact paths. A training preference is a content-policy decision, not a reason to block search access by default.

Scope the disallow

Decide whether the policy covers the whole site or selected directories. A complete block is easy to explain but may affect public pages you wanted in search. A path-specific rule needs careful testing on both included and excluded URLs.

Know the limits

Robots.txt communicates a requested crawl policy. It is not a legal contract, a firewall rule or a deletion request, and provider handling differs. Review your terms, privacy policy and provider documentation with the right adviser for your organisation.

Test the live file

After deployment, fetch robots.txt from the public hostname and run a check for the named crawler. Keep the old and new files in the change record. If a CDN or WAF is involved, review its controls separately.

Read the current bot policy

Use the OpenAI bot documentation for the named crawler and date the review. Reviewed 7 October 2026.

Questions

Does blocking GPTBot block OAI-SearchBot?

Not automatically. Check both named groups because they can represent different purposes.

Does a robots block remove content already collected?

No. A new rule does not undo earlier fetching or copies. Treat it as a crawl policy for future requests.

Related resources