Firm Beacon

AI crawler access explained: robots.txt, search and training

AI crawler access means robots.txt permits a named user agent to request a path. It does not prove a visit, indexing or an AI answer. Check search and training user agents separately.

Updated 7 October 2026. Written for website owners and teams in the US and Europe.

Direct answer

AI crawler access means robots.txt permits a named user agent to request a path. It does not prove a visit, indexing or an AI answer. Check search and training user agents separately.

Robots.txt is a matching rule set

A robots.txt file contains user-agent groups and path rules. The relevant crawler applies the group that matches it, then follows the path rules under that group. A wildcard group is not the same as a provider-specific group when a more specific group exists. Matching path length and Allow rules can affect the result.

The exact semantics matter when a site has several groups, inherited assumptions or a private directory near a public resource. Inspect the file that the public URL returns, not a screenshot or a cached fragment.

Search and training are different decisions

A provider may use different crawler identities for search access and model training. OpenAI documents OAI-SearchBot for ChatGPT search and GPTBot for training. A policy can allow one and block the other. The names and policies can change, so link to current provider documentation when a decision matters.

Use current provider references

For the current OpenAI bot names and purposes, check the OpenAI bot documentation. For rule matching and group behavior, use Google's robots.txt documentation. These links were checked on 7 October 2026; keep the review date with any production policy change.

What the check cannot prove

A robots result cannot prove that a crawler made a request, that a CDN or firewall allowed it, that a page is indexed, that an answer system selected it or that a human saw a citation. It also does not inspect page content, meta noindex tags or access controls outside robots.txt.

State those limits in an internal ticket and in any public explanation. A clear boundary prevents a team from treating a permission file as an analytics report.

Check the exact path

A site may allow its homepage while disallowing a product directory or an article path. Enter the URL that matters. Then check a second URL from a different directory when the policy is complex. Keep private areas blocked and test a proposed rule before publishing it.

When an allowed crawler still cannot fetch

Check DNS, TLS, authentication, a CDN rule, a WAF challenge and the origin response. A copied user-agent string from a browser is not proof of a genuine provider crawler. Work with access logs or provider verification when a high-stakes decision depends on the request.

A sensible change record

Record the old rule, the new rule, the affected path, the intended crawler and the reason for the change. After publishing, fetch robots.txt again and run the rule check on the exact page. Keep the date because provider user agents and site policies evolve.

Use the result as one part of readiness

Pair crawler access with a technical HTML audit, a clean sitemap and clear page content. Those checks answer different questions. None guarantees an AI answer, but together they make the site's public behavior easier to inspect and maintain.

International scope

The rule interpretation is not tied to one country's legal system. It is written for website owners and teams in the US and Europe. Provider documentation, privacy rules and internal content policies still require local review where they apply.

Policy examples, not universal templates

A site that publishes product documentation may allow search crawlers across its public documentation while blocking internal previews. A publisher that does not want training requests may express a separate policy for the relevant training user agent. A private customer portal may block both. These examples illustrate decisions, not a universal robots.txt file.

Write the policy in plain language before editing the file: which provider, which purpose, which paths and which exception? Then map that sentence to a user-agent group and test it. If the mapping is unclear, do not add a rule just because a snippet appears in a blog post. Keep the source of the policy and the test URL in the change record.

Review the result with the people who own content, infrastructure and privacy. The technical check can show that a path matches a rule, but it cannot decide whether the organisation should make that path available. That decision belongs in the policy, with the crawler check used as evidence that the live file expresses it.

When the file is not enough

A robots review should trigger a second question when the site uses a CDN, WAF, login, geo restriction or a custom crawler policy. Ask which system can reject the request after robots.txt has allowed it, and who can provide evidence of that rejection. This prevents a team from endlessly editing a text file to solve an infrastructure problem. Keep the public rule result and the server-side evidence in separate fields.

Keep human decisions visible

Access policies often affect more than search. A content owner may want a public article discoverable, while a privacy or licensing policy may restrict a directory. Put the decision and its owner next to the technical rule. A crawler check can show that the live file matches the decision, but it cannot choose the acceptable tradeoff for the organisation.

Review after provider changes

When a provider changes a crawler name or purpose, revisit the policy and the test cases. Keep old observations for context, but do not assume a result from last year describes today's user agent. A dated, small review is enough to expose a mismatch before it becomes a publishing surprise.

Questions

Does allowing OAI-SearchBot guarantee ChatGPT traffic?

No. It removes a possible robots barrier. It does not prove fetching, indexing, selection or traffic.

Does blocking GPTBot block ChatGPT search?

Not necessarily. OAI-SearchBot and GPTBot are separate user-agent groups, so check each one.

Can robots.txt block every AI crawler?

It can express rules for named user agents, but enforcement, verification and provider behavior are separate concerns.

Related resources