1. Technical SEO Foundation (11 Points)
Traditional technical SEO forms the prerequisite layer for any website. If search engines or LLM crawlers encounter server timeouts, broken canonicals, or invalid HTTP responses, downstream AI search optimization cannot function.
- 1.1 HTTPS Encryption: Enforce valid SSL/TLS encryption across all routes with 301 redirects from HTTP to HTTPS.
- 1.2 Indexability & Meta Robots: Verify that public pages do not contain accidental
noindexornofollowdirectives. - 1.3 Valid robots.txt: Maintain a syntactically correct
robots.txtfile at your domain root with a clearSitemap:reference. - 1.4 XML Sitemap Health: Ensure
/sitemap.xmlcontains canonical 200 HTTP status URLs and is submitted to search consoles. - 1.5 Canonical URL Consistency: Ensure every page specifies an absolute
rel="canonical"URL matching the preferred domain version. - 1.6 Clean HTTP Status Codes: Eliminate 404 Not Found errors and 5xx server exceptions on core content pages.
- 1.7 Minimal Redirect Chains: Avoid multi-hop 301/302 redirect chains that waste crawler budget.
- 1.8 JavaScript SSR Rendering: Ensure critical text content and metadata are rendered server-side so bots without heavy JS execution can extract content.
- 1.9 Mobile Usability: Ensure viewport tags are configured correctly (
width=device-width) with responsive CSS layout touch targets. - 1.10 Core Web Vitals (LCP, CLS, INP): Optimize Largest Contentful Paint (< 2.5s), Cumulative Layout Shift (< 0.1), and Interaction to Next Paint (< 200ms).
- 1.11 Server Speed & TTFB: Maintain Time to First Byte under 800ms via CDN caching and optimized database queries.
2. AI Crawler Accessibility (8 Points)
Many site owners unknowingly block AI search crawlers in their robots.txt file while attempting to block training data scrapers. An AI SEO audit explicitly verifies crawler permissions for conversational agents.
- 2.1 AI User-Agent Permissions: Verify rules for
GPTBot,ChatGPT-User,PerplexityBot,ClaudeBot, andGoogle-Extendedinrobots.txt. - 2.2 High-Value Content Access: Ensure documentation, pricing pages, and product guides are accessible to AI user-agents.
- 2.3 Consistent Server Responses: Confirm that requests with AI user-agent headers receive standard 200 OK responses without firewall CAPTCHA blocks.
- 2.4 Unblocked Asset Dependencies: Ensure CSS stylesheets and static JSON API endpoints are not blocked by crawler disallow rules.
- 2.5 Client-Only Hydration Safety: Verify that content is not hidden behind user click actions or unrendered client-side React state wrappers.
- 2.6 Clean HTML Text Extraction: Ensure primary article text is contained inside standard semantic HTML elements (
<p>,<h1>-<h3>,<article>). - 2.7 Sitemap Discovery for AI Indexers: Provide a direct link to XML sitemaps inside
robots.txtfor fast AI crawler discovery. - 2.8 AI Crawler Rate Limits: Ensure web hosting server firewalls do not rate-limit legitimate search agent user-agents.
3. Content & Entity Clarity (9 Points)
AI search engines use vector embeddings to match user queries with web content. Ambiguous, unfocused, or fluff-filled articles fail semantic vector reranking.
- 3.1 Single Core Topic Focus: Each page should address one primary subject with clear thematic boundaries.
- 3.2 Search Intent Alignment: Match informational vs. transactional search intent without forced keyword padding.
- 3.3 Factual & Verifiable Claims: Ensure technical statements and statistics are accurate and cross-verifiable.
- 3.4 E-E-A-T Author Credentials: Display clear author names, publisher credentials, and publication dates.
- 3.5 Entity Relationship Clarity: Connect brand names, key concepts, and product names with clear semantic syntax.
- 3.6 Topical Depth: Provide comprehensive coverage of core sub-topics rather than shallow repetitive paragraphs.
- 3.7 Original Insights & Data: Include first-party data, original charts, or unique analysis that cannot be duplicated by simple LLM training data.
- 3.8 Real Product/Service Examples: Illustrate concepts with verified, concrete real-world use cases.
- 3.9 Unambiguous Answers: State direct answers immediately following question subheadings.
4. Answer Engine Optimization / AEO (6 Points)
Answer Engine Optimization (AEO) prepares content for direct snippet extraction in voice search and conversational AI answers.
- 4.1 Direct Answer Summary Boxes: Place concise 2-sentence summary boxes at the top of key article sections.
- 4.2 Question-Based Headings: Format H2/H3 subheadings as natural language questions (e.g.,
What is AEO?). - 4.3 Concise Definition Snippets: Start answer paragraphs with a clear, standalone subject definition.
- 4.4 Visible FAQ Sections: Include visible FAQ sections matching schema markup.
- 4.5 Structured Tables & Lists: Present technical comparisons using clean HTML
<table>and<ul>tags. - 4.6 Conversational Search Intent: Target natural voice-search phrasing alongside technical keyword terms.
5. Generative Engine Optimization / GEO (6 Points)
Generative Engine Optimization (GEO) focuses on building entity authority to increase the likelihood that generative AI search overviews cite your brand as an authoritative source.
- 5.1 Factual Content Density: Maintain high factual density with concrete measurements and zero generic fluff.
- 5.2 Authoritative Source References: Cite official documentation, W3C standards, and technical specifications.
- 5.3 Brand Entity Disambiguation: Use consistent brand name syntax across all digital footprints.
- 5.4 Original Research Citations: Publish transparent methodology studies that invite third-party citations.
- 5.5 Verifiable Statistics: Include cited numbers and date timestamps.
- 5.6 Summarizable Passage Architecture: Structure long paragraphs into self-contained 3-sentence passages.
6. Structured Data & Schema.org Realities
Structured data (JSON-LD) builds entity clarity for machine indexers. However, it is essential to understand that schema markup provides machine readability, not guaranteed search rankings or automated AI citations.
Organization Schema
Defines brand name, logo URL, official website, and official social media profiles (sameAs).
WebSite Schema
Establishes domain identity and search query action templates (SearchAction).
WebApplication / SoftwareApplication
Describes software tool features, operating systems, and price offers (e.g., price: "0").
Article & BreadcrumbList Schema
Identifies editorial content, publication dates, author credentials, and navigational hierarchy.
FAQPage schema when the questions and answers are genuinely visible in the page's HTML body. Do not embed hidden FAQ schema scripts.7. The llms.txt Context Standard
llms.txt is an emerging web standard designed to serve lightweight markdown summaries of website documentation and key pages directly to LLM crawlers.
- Purpose: Provides a clean, unstyled markdown index of a website's core pages, APIs, and product offerings so LLM context windows do not waste token capacity parsing heavy HTML/JS templates.
- Implementation: Placed at your domain root (e.g.,
https://aiscanmysite.com/llms.txt) following standardized markdown H1/H2 syntax. - Limitations:
llms.txtis an informational context feed—it does not overriderobots.txtblock rules or guarantee indexing. - Validation: Test markdown formatting to ensure links are absolute and syntactically clean. Generate your manifest using our free llms.txt Generator.
llms.txt is a voluntary opt-in informational feed. It does not guarantee that ChatGPT, Claude, Perplexity, Gemini, or other AI platforms will cite, index, or link to your website. It is not an automated ranking signal or citation mechanism. Treat it as documentation for LLM context optimization, not a substitute for content quality, backlinks, or traditional SEO.
8. The agents.json Protocol
While llms.txt provides readable text context, agents.json serves a distinct role: it defines operational action manifests for autonomous AI agents.
Key Differences Between llms.txt & agents.json
- llms.txt (Markdown Context Feed): Formatted for LLM text understanding, summarizing page titles, documentation links, and brand descriptions.
- agents.json (Machine Action Manifest): Formatted in structured JSON to define API endpoints, authentication mechanisms, and functional capabilities for autonomous AI agents performing transactions.
Audit Your Website Now for Free
Don't guess whether your site is accessible to AI search engines. Run a free real-time audit using AI Scan My Site to inspect robots.txt rules, llms.txt manifests, JSON-LD schemas, and PageSpeed Core Web Vitals in 10 seconds.
Run Your Free AI SEO Audit
Instant diagnostic report for ChatGPT, Perplexity, Gemini, and Google search readiness.