AI SEO technical checklist
Before trying to improve GEO or AI search visibility, remove the technical barriers first. This checklist separates what you can verify from what no tool can promise.
The short version
AI SEO is not one special meta tag. Start with the same foundation as ordinary SEO: pages must be reachable, understandable and worth referring to. Then check the crawler rules for the specific AI search services you want to reach.
- Choose the exact pages you want discovered, not just the home page.
- Check robots.txt for the relevant search crawler and keep training policies separate.
- Check noindex, canonical tags, login walls, redirects and firewall rules independently.
- State facts plainly, identify the organisation or author, and link important claims to primary sources.
- Measure real crawl requests, search impressions and referral traffic. Do not treat a robots.txt result as a visibility score.
1. Check access to the exact page
A site can allow its home page and still block a product page, article or documentation path. Test the URL that contains the answer you want people to find.
Our free AI crawler access checker reads the matching robots.txt rules for OAI-SearchBot, GPTBot, Claude, Perplexity, Google and Bing. It can fetch a public robots.txt file or analyse pasted rules without an account.
Use separate groups when your policy differs. For example, allowing a search crawler while blocking a training crawler is a different decision from allowing both:
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
This example only describes robots.txt permission. It does not prove that a CDN, firewall, origin server or provider will fetch the page.
2. Look for barriers robots.txt does not cover
Use the free website audit to review a sample of public pages, or follow the SEO audit workflow for indexing and performance checks that require separate tools.
- Meta robots and headers. Check whether the page or response headers contain noindex or other directives that conflict with your goal.
- Canonical URLs. Make sure the canonical points to the version you want discovered and that it is not redirected to an unrelated page.
- Authentication. A page hidden behind a login cannot be read by a public crawler.
- CDN and firewall rules. Edge challenges or bot policies can stop a request before it reaches your origin.
- Rendering. Put essential facts in the delivered page, not only in a client-side interaction that a crawler may not execute.
Keep these findings separate. A robots.txt permission is not evidence of a successful fetch, indexing, citation or traffic.
3. Make the answer easy to extract and check
Write the answer near the start of the page. Use a descriptive heading, short paragraphs, useful tables or lists, and stable links to supporting pages. Explain terms before using an abbreviation such as GEO or AEO.
For factual or commercial pages, make the source of important claims clear. Include the organisation name, the relevant date, the scope of the claim and a link to the primary document when one exists. Do not add FAQ markup to content that is not visible on the page.
Structured data can help systems interpret a page, but it cannot make unsupported claims true. Keep the visible wording and structured data consistent.
4. What to measure
- Search Console impressions and clicks for the pages you changed.
- Server or CDN logs showing verified crawler requests, where the provider publishes a way to identify them.
- Referrals from AI services, recorded separately from ordinary search traffic.
- Whether the page is indexed and whether real users can reach it without a challenge.
A tool that checks robots.txt can answer one narrow question: which rule matches a crawler and path. It cannot tell you whether an AI system will cite the page or send visitors.
Common GEO mistakes
- Blocking all crawlers when the actual goal was only to block training.
- Assuming that an Allow rule overrides a CDN or firewall challenge.
- Publishing generic AI-written pages with no original facts, sources or reason to trust them.
- Calling a robots.txt check an AI visibility audit.
- Measuring success from a crawler hit instead of from indexing, impressions or qualified visits.
Sources
OpenAI's crawler documentation describes separate crawler purposes. The Google robots.txt guide explains user-agent groups and path rules. Check each provider's current documentation before changing access policies.
This checklist is technical guidance, not a guarantee of search visibility or legal advice. Crawler policies and search systems can change.