Firm Beacon

How to allow AI search crawlers without opening private paths

To allow an AI search crawler safely, inspect the user-agent groups, add the narrowest matching rule, publish it, and test the live robots.txt response. Keep private paths disallowed.

Updated 7 October 2026. Keep the output and the verification step together.

Steps

  1. List the public pages you want discovered and the private paths that must remain closed.
  2. Check the current groups for OAI-SearchBot, GPTBot and the wildcard user agent.
  3. Choose whether search access and training access should have different policies.
  4. Edit the smallest group and path that expresses the decision; do not delete existing privacy rules.
  5. Fetch robots.txt from the live hostname and run the check on one public URL and one private URL.
  6. Record the date and review provider documentation again when the policy changes.

Why narrow rules matter

A broad Allow: / can expose more than the page you intended. A provider-specific group can also change how wildcard rules apply. Read the file as a set of matching groups, not as a list of independent lines.

Allow one market's public path

For an international site, start with one public path such as `/en-us/docs/` and one private path such as `/en-us/internal/`. Test the same pair on the European path before widening the rule. This keeps an allow decision tied to a release scope instead of opening every regional directory at once.

What to verify after publishing

Check the live file, the exact path and the relevant user-agent. Then inspect access logs or provider verification if the request itself matters. Robots.txt alone cannot show whether a crawler visited.

Keep the policy readable

Add a short comment for the team if your server supports it, and store the source file with the deployment. Future editors should understand why search and training groups differ.

Use current references

Check the OpenAI bot documentation before naming a search crawler and the Google robots.txt documentation before relying on matching behavior. Reviewed 7 October 2026.

Questions

Can I allow search and block training?

Often, yes, when the provider uses separate crawler names. Check current provider documentation and your own policy.

Does an Allow rule override every block?

Matching path length and the crawler's selected group matter. Test the actual file rather than relying on a slogan.

Related resources