---
title: "GEO & Agentic Checklist: SI Readiness Guide | Webtures"
description: "An actionable checklist for generative engine optimisation and agent readiness, covering content, technical foundations, structured data and SI platform"
source_url: "https://www.webtures.com/geo-checklist/"
lang: "en"
---

# GEO & Agentic Checklist

An actionable checklist for generative engine optimisation and agent readiness, covering content, technical foundations, structured data and SI platform visibility.

## GEO and Agent-Readiness Checklist Sections

Generative Engine Optimization

Progress 0 / 164 completed

Hide completed

Content Comprehensibility and Quality

0/15

Language and Expression Write in clear, plain language with unambiguous sentences so SI engines parse meaning accurately and reuse your wording in generated answers. High

Understandability for AI Make information machine-graspable, not just human-readable: explicit context, resolved references and self-contained statements that survive extraction. High

Headings and Subheadings Use descriptive, question-shaped headings that map to real queries so SI systems locate and lift the matching passage. Medium

Content Longevity Maintain evergreen, regularly refreshed content; SI engines favour durable sources and cite freshly updated pages more often. Medium

Recommendations and Citations Back claims with sources and outbound citations so SI engines trust the passage and are likelier to reference it. High

List Content Present steps and key points as bullet or numbered lists, a structured format SI systems parse and reproduce reliably in answers. High

Date Indication Show clear publish and update dates so SI engines judge freshness and prefer your page for time-sensitive answers. Medium

Brand and Authority Reinforce author and brand authority signals throughout so SI engines weigh your content as a credible, citable source. Medium

Direct Answer Summaries Open each section with a self-contained 40-60 word answer to the implied question so SI systems can lift it as a citation. High

Key Takeaways Sections Add a concise takeaways block to each article so models can extract the core points without parsing the full body. High

Definition Boxes for Key Terms Provide clear, standalone definitions for important terms so SI engines can quote precise, unambiguous explanations. Medium

Comparison Tables Present options, features or metrics in structured tables, a format SI systems parse and reproduce reliably in answers. Medium

Quotable Statistics and Data Points Include concrete figures, percentages and dated data points, which SI engines cite far more often than generic claims. High

Expert Quotes and Attribution Add named expert quotes with attribution to raise authority and give models a clear, citable source. Medium

In-Content FAQ Sections Answer real follow-up questions in question-and-answer blocks within the content, beyond FAQPage schema, to match conversational queries. Medium

Technical Optimization

0/7

Page Speed Keep load times and Core Web Vitals fast so SI crawlers fully render and index your content without timing out. High

HTTPS Usage Serve every page over HTTPS; SI crawlers and answer engines favour secure, trustworthy sources. High

Structured Data (Schema Markup) Add schema markup so SI engines read entities, relationships and facts machine-readably and cite them confidently. Medium

Crawling and Indexing Ensure key pages are crawlable and indexed; content SI engines cannot reach cannot be cited. Medium

IndexNow Protocol Ping search engines such as Bing and Yandex via IndexNow so new or updated pages are crawled in minutes instead of days, keeping SI retrieval fresh. Medium

Broken Links and Redirect Chains Eliminate 404s and collapse multi-hop 301 chains so crawlers and SI agents reach final content without wasted budget or dead ends. High

Snippet and AI Preview Controls Set max-snippet, max-image-preview and data-nosnippet directives deliberately so SI features can surface the passages you want cited. High

User Experience and Interaction

0/1

User Comments and Interaction Enable genuine user comments and interaction; engagement and fresh user content add signals SI engines read for relevance and trust. High

Social Signals and Backlinks

0/4

Social Media Shares Earn social shares that amplify reach and discovery, indirect signals that help SI engines find and corroborate your content. High

Backlinks Build relevant inbound links; authoritative backlinks remain a core trust signal SI engines inherit from search. Medium

Social Media Integration Connect and reference your social profiles so SI engines link them as part of a consistent brand entity. Medium

Backlink Quality Prioritise links from high-authority, topically relevant domains; SI engines cite pages backed by strong link profiles far more often. Medium

Structured Data

0/14

JSON-LD Format Use JSON-LD for structured data, the format SI engines and search parse most reliably for entity extraction. High

Organization Schema Describe your organisation with Organization schema so SI engines anchor your brand as a recognised entity. High

FAQPage Schema Mark up questions and answers with FAQPage schema so SI engines extract direct answers to conversational queries. High

Product or Service Schema Add Product or Service schema so SI engines surface your offerings with accurate, structured detail. High

BreadcrumbList Schema Provide BreadcrumbList schema so SI engines understand site hierarchy and page context. Medium

Review and AggregateRating Schema Expose ratings with Review and AggregateRating schema, social-proof signals SI engines summarise in recommendations. High

Article Schema (author, datePublished, dateModified) Mark articles with author, datePublished and dateModified so SI engines attribute authorship and assess freshness. Medium

Speakable Schema (for voice answers) Add Speakable schema to flag passages suited to spoken answers from SI voice assistants. Medium

sameAs Property (Wikidata, LinkedIn, Crunchbase) Link your entity to Wikidata, LinkedIn and Crunchbase via sameAs so SI knowledge graphs reconcile your identity. High

Critical Properties (@type, name, url) Always include @type, name and url, the core properties SI engines need to resolve an entity. High

Rich Properties (offers, availability, priceRange) Add rich properties such as offers, availability and priceRange so SI engines present complete, accurate detail. Medium

Person or Author Schema Mark up authors as Person entities with credentials and sameAs links so SI engines attribute expertise and strengthen E-E-A-T signals. Medium

HowTo Schema (instructional content) Add HowTo structured data to step-by-step guides so answer engines can extract and present the procedure directly. Medium

Structured Data Validation Validate every schema block with the Rich Results Test and Schema.org validator to catch errors that silently block eligibility. High

SI Discoverability

0/14

llms.txt File Publish an llms.txt file pointing SI models to your most important content, raising machine readability and citation odds. High

llms-full.txt File Provide an expanded llms-full.txt with deeper content so SI models access full context, not just summaries. Medium

llms.txt Content Quality Keep llms.txt accurate and well curated so the signals you send SI models reflect your strongest content. Medium

AGENTS.md File Add an AGENTS.md file to guide SI agents on how to use your site and content. Medium

skill.md File Provide a skill.md describing capabilities SI agents can invoke, aiding agent-driven discovery. Low

AI Bot Access in robots.txt Explicitly allow GPTBot, PerplexityBot, ClaudeBot and Google-Extended in robots.txt; to be cited, you must be crawlable. High

Canonical URLs Set canonical URLs so SI engines consolidate signals on one authoritative version of each page. High

Meta Descriptions Write precise meta descriptions; SI engines and previews draw on them to summarise and surface your pages. High

Open Graph Tags Add Open Graph tags so SI engines and platforms render accurate titles, descriptions and images when referencing you. Medium

XML Sitemap Submit an XML sitemap so SI crawlers discover every important page efficiently. High

Sitemap Freshness (lastmod) Keep lastmod values accurate so SI crawlers prioritise recently updated content. Medium

Mobile Viewport Meta Declare a mobile viewport so SI engines render and assess your pages as mobile-friendly. High

Low Error Rate (no 4xx/5xx) Keep 4xx and 5xx errors near zero so SI crawlers reach content reliably and your authority is not diluted. High

ai.txt File Publish an ai.txt file to declare SI usage and access preferences alongside llms.txt and robots.txt for emerging crawler conventions. Low

Content Accessibility

0/6

Heading Hierarchy (H1, H2, H3) Maintain a logical H1-H3 hierarchy so SI engines map document structure and locate the right passage to cite. High

Image Alt Text Add descriptive alt text so SI engines understand visual content and can reference it in multimodal answers. High

Semantic HTML (article, section, nav, aside) Use semantic elements such as article, section, nav and aside so SI engines interpret content roles and boundaries accurately. Medium

Content Structure (intro, body, conclusion) Follow a clear intro, body and conclusion structure so SI engines extract coherent, self-contained passages. Medium

Internal Linking Connect related pages with descriptive anchor text so crawlers and SI agents discover topical depth and understand entity relationships. High

Thin or Duplicate Content Consolidation Remove or merge thin and duplicate pages so authority concentrates on the canonical content SI engines prefer to cite. High

API and Agent Protocols

0/6

API Endpoint (JSON or REST) Expose a JSON or REST endpoint so SI agents can access your data programmatically as a structured source. Medium

CORS Permissions Configure CORS so SI agents and tools can fetch your resources across origins without being blocked. Low

OpenAPI (Swagger) Documentation Document APIs with OpenAPI so SI agents understand and call your endpoints correctly. Medium

MCP Manifest (Model Context Protocol) Provide an MCP manifest so SI assistants can connect to your data and actions as a context source. Medium

Function Calling Ready Expose well-described functions so SI models can invoke your capabilities through function calling. Medium

OpenAI Plugin or Actions Definition Define a plugin or Actions schema so SI assistants can integrate your service directly into their answers. Low

SI Platform Visibility

0/10

Citation in ChatGPT Track whether ChatGPT cites you for core queries, a primary indicator of GEO success on the largest SI platform. High

Google AI Overview and AI Mode Visibility Monitor appearances in Google AI Overviews and AI Mode, where SI answers increasingly replace classic results. High

Perplexity Source Listing Check whether Perplexity lists you as a source; it surfaces citations prominently and rewards fresh, authoritative pages. High

Citation in Claude and Gemini Verify citations in Claude and Gemini to confirm visibility across multiple SI ecosystems, not just one. Medium

Bing Copilot Visibility Track Copilot visibility; the Bing index also feeds ChatGPT retrieval, making it doubly important. Medium

Weekly Citation Tracking (prompt corpus) Run a fixed prompt corpus weekly to measure citation frequency, position and share of voice over time. High

AI Referral Traffic Tracking (GA4) Define SI sources as a dedicated channel in GA4 to measure sessions and conversions arriving from ChatGPT, Perplexity and similar engines. High

AI Overviews Monitoring in Search Console Watch impressions and query shifts in Google Search Console to infer where AI Overviews are surfacing your pages. Medium

Competitor Citation Monitoring Track which rival brands SI engines cite for your core queries to expose share-of-model gaps you need to close. Medium

AI Search Performance Dashboard Consolidate citation frequency, sentiment and referral metrics into one dashboard for ongoing, comparable reporting. Medium

Content Strategy and Topic Coverage

0/7

Target Keyword AI Visibility Audit Test your priority queries across ChatGPT, Gemini and Perplexity to baseline where you appear and how you are described. High

Content Gap Analysis Map the questions SI answers for your category that your content does not yet cover, then prioritise filling them. High

Pillar and Cluster Content Architecture Build topic pillars linked to supporting cluster pages so SI engines recognise depth and topical authority. High

Ultimate Guide Content for Core Topics Publish comprehensive, definitive guides on core topics, the long-form format SI systems favour as authoritative sources. High

Conversational and Long-Tail Optimization Optimise for natural-language, question-shaped and long-tail queries that mirror how users prompt SI assistants. High

Original Research and Data Produce proprietary studies, surveys and data sets that become uniquely citable and earn references across SI platforms. High

Multi-Format Content Offer the same topic in text, visual, video and infographic formats to widen retrieval surfaces across engines. Medium

Brand Entity and E-E-A-T

0/8

Brand Information Consistency Ensure name, description, offerings and contact details match across every property so SI engines form one coherent entity. High

Google Business Profile Create and optimise a Google Business Profile to strengthen entity signals and local SI visibility. High

Wikipedia and Wikidata Presence Establish accurate Wikipedia and Wikidata entries, primary sources that anchor brand identity in SI knowledge graphs. High

Comprehensive Author Pages Build detailed author pages with bios, credentials and sameAs links so models can verify who stands behind the content. High

Comprehensive About Page Provide a thorough About page covering history, expertise and team, a high-signal source SI engines read for trust. High

Executive and Team Thought Leadership Publish named expert perspectives so individuals become recognised entities that reinforce brand authority. Medium

LinkedIn Profile Optimization Keep key personnel profiles complete and consistent so SI engines corroborate identity and expertise off-site. Medium

Industry Directory Listings Claim and align listings in reputable industry directories to reinforce consistent entity data across the web. Medium

Off-Site Reputation and Authority

0/5

Brand Mention Monitoring Track brand mentions across forums, reviews and communities, signals SI engines weigh heavily when forming opinions. High

Industry Discussion Participation Contribute genuinely on community platforms such as Reddit and Quora, frequent sources for SI citations. Medium

Customer Review Generation Encourage detailed reviews on trusted platforms to build the social proof SI systems summarise and cite. Medium

Publishing on High-Authority Platforms Earn placements and contributions on high-authority external sites that SI engines crawl and trust. Medium

Negative Mention Response Respond to negative mentions professionally so the sentiment SI engines synthesise stays balanced and fair. Medium

Competitive Intelligence

0/4

Citation Share of Voice Benchmark Measure how often SI engines cite you versus category rivals to quantify your share of model answers. High

Cited Content Analysis Study the rival content SI engines cite to identify the formats, depth and signals that earn references. High

Entity Strength Analysis Assess where competing entities are stronger in SI knowledge graphs to target your own entity gaps. Medium

Emerging AI Platform Tracking Monitor new SI search and answer platforms early to capture visibility before competition intensifies. Medium

Crawlability and Server Health

0/12

Readable Without JavaScript Verify the content is readable with CSS, JavaScript and cookies disabled; most SI crawlers read raw HTML and never see client-rendered text. High

Navigation Works Without JavaScript Check that menu and navigation links work without JavaScript; otherwise the discovery path to your pages closes. High

No Server Errors Find and fix pages returning 500 status codes; erroring pages burn crawl budget and drop out of the index. High

Indexed 404 Pages Clear indexed 404 pages; SI engines will not treat them as sources and they weaken the site's quality signal. Medium

Unintended 302 Redirects Make sure 302 is not used where a move is permanent; a temporary redirect leaves authority transfer ambiguous. Medium

Single-Hop 301 Redirects Verify every 301 reaches its final 200 URL in one hop; chains raise crawl cost and risk losing the citation link. High

Meta Refresh Redirects Replace meta refresh redirects with server-side 301s; crawlers do not treat them as a reliable signal. Low

Externally Sourced iframe Content Ensure iframe content also exists in the page's own HTML; iframe content does not count as a citable passage. Medium

Text Embedded in Images Deliver text as real HTML instead of baking it into images; text inside an image is unreadable to SI engines. Medium

Cloaking Check Confirm bots and users receive the same content; when detected, you are removed from both the classic index and SI surfaces. High

Blocking Non-Content Pages Keep cart, account, print and filter pages out of crawling so crawl budget goes to citable content. Medium

Drops in Crawl Stats Watch for sudden drops in Search Console crawl stats; a drop is usually the earliest warning of lost visibility on SI surfaces. Medium

Indexability and Duplication

0/12

Know Your Indexed Page Count Measure how many pages are indexed; AI Overviews and AI Mode are grounded in the same index, so a page that is not indexed is not in answers either. High

First Position on Brand Queries Your own site must rank first for your brand name; it is the most basic indicator of entity clarity and brand authority. High

Indexing of JavaScript-Rendered Content If content loads via JavaScript, confirm it is also produced server-side; server rendering is the safest route to SI access. High

Count of Uncrawlable Pages Measure how many pages bots cannot reach and remove the cause; every unreachable page is a lost citation opportunity. Medium

No Full Copy of the Site Check that the same content is not published on a second domain; a duplicate entity makes attribution ambiguous. High

Protocol and www Duplication Verify HTTP/HTTPS, trailing slash and www variants resolve to one version; split signals lower your odds of being cited. High

Correct Canonical Tags Audit that canonicals point to the right page; a wrong canonical sends the citation to the wrong address. High

Duplicate Pages, Titles and Descriptions Consolidate duplicate pages, titles and meta descriptions; repetition weakens an engine's source selection. Medium

Sitemaps by Content Type Produce separate sitemaps for web, image, video and news content; discovery speed and scope clarity both improve. Medium

Every Sitemap URL Returns 200 Verify each URL in the sitemap returns 200; a map containing 301s or 404s lowers the trust signal. High

Sitemap Submitted to Search Console Submit the map through Search Console and track the error report; it surfaces discovery problems fastest. Medium

URL Parameters Kept Out of the Index Check that filter and tracking parameter URLs are not indexed; parameterised copies split authority. Medium

Performance and Core Web Vitals

0/10

Resolve Core Web Vitals Issues Track and fix LCP, INP and CLS through Search Console; page experience is a hygiene condition for both classic ranking and source selection. High

Server Response Time Reduce server response time; slow-responding pages are the first to have their crawl frequency cut. High

Image Optimisation and WebP Serve images in modern formats such as WebP at the right dimensions; page weight directly affects crawlability. Medium

Defer Offscreen Images Lazy-load images below the fold; first paint arrives sooner and content becomes accessible earlier. Medium

Minify JS and CSS Minify and bundle JavaScript and CSS, and avoid heavy use of inline code. Medium

Text Compression Confirm text compression is enabled on the server; smaller transfers mean lower crawl cost. Medium

Caching Check that static assets are cached correctly to speed up repeat requests. Medium

CDN Usage Distribute content through a CDN; geographic latency drops for users and bots alike. Medium

DOM Size Keep the DOM node count reasonable; an oversized DOM slows both rendering and content parsing. Low

Font Loading Cost Avoid heavy, many-variant font stacks; text becomes visible sooner. Low

Site Architecture and Internal Link Depth

0/9

Clear Category Structure Give every category and subcategory a clear purpose and a sensible depth; a vague hierarchy scatters the topical authority signal. High

URLs Reflect the Hierarchy Make sure the URL reflects where the page sits in the site; the address alone should convey the content's context. Medium

Priority Pages Near the Surface Keep commercially important pages a few clicks from the homepage; deeply buried pages are crawled less and cited less. High

Breadcrumbs Build breadcrumbs and mark them up with BreadcrumbList schema; it tells the machine exactly what context to read the page in. Medium

Followable Main Navigation Ensure top and subcategory links are reachable from the homepage and marked dofollow. Medium

Internal Linking Strategy Build a deliberate internal linking plan that connects topic clusters; SI engines relate passages through those links. High

Finding Orphan Pages Identify pages that appear only in the sitemap with no incoming links, and link them. Medium

One Address for the Homepage Confirm the homepage is not reachable from several addresses; duplication splits the signal of your strongest page. Medium

Pagination Markup On paginated lists, check links point to the right addresses and that canonicals do not override pagination. Low

Measurement Foundations

0/7

Search Console Verification Verify the site in Search Console; the generative search performance report is only readable there. High

Bing Webmaster Tools Verification Complete Bing Webmaster Tools registration; it is the only first-party source for Copilot visibility. Medium

Pages Missing Analytics Code Find and fix pages without measurement code; missing measurement makes SI-sourced traffic invisible. High

Site Search Tracking Measure on-site search queries; what visitors type in their own words is your rawest prompt source. Medium

UTM and Campaign Parameters Use campaign parameters consistently; it is a precondition for separating traffic arriving from SI surfaces. Medium

Event Tracking and Goals Set events with the right category and action and mark them as goals; this is the only way to see high-quality traffic that arrives in low session counts. Medium

Data Sampling Limits Check reports are not built on heavily sampled data; low-volume SI traffic disappears under sampling. Low

Product and Commerce Data

0/7

A Product Feed Exists Publish a current, complete product feed; agent-driven shopping flows read product data from it first. High

Complete Feed Fields Verify every required feed field is filled; a missing field leaves the product out of comparisons. High

Images in the Feed Ensure feed records include images; records without one are not shown in multimodal answers. Medium

Feed Update Frequency Refresh price and stock data regularly; wrong stock data is the most common failure in agent-driven purchases. High

Unique Product Descriptions Write original product copy instead of reusing the manufacturer's text; distinctive text decides your chance of being cited. High

Canonicals on Variants If variants sit on separate addresses, canonicalise them to the main product so authority consolidates on one page. Medium

Duplication in Faceted Navigation If faceted navigation is in use, audit the duplicate addresses it generates; filter combinations inflate the index quickly. Medium

Mobile and Device Parity

0/6

Mobile and Desktop Content Parity Confirm no content is hidden on mobile; indexing runs on the mobile version, so text missing there is treated as absent. High

Unblock Resources Check JavaScript, CSS and image resources are not blocked on mobile; a blocked resource leads to the page being misread. High

Mobile Page Speed Measure mobile speed separately; a page that performs well on desktop can fall below the crawl threshold on mobile. High

Mobile Readability Set body text at 16px or larger with comfortable line height on mobile; readability feeds engagement signals directly. Medium

Separate Mobile Addresses If a separate mobile address is in use, verify the rel=alternate and canonical pairing, and move to a single responsive address where possible. Low

Conversational Queries Fold the full-sentence queries that appear in mobile and voice use into the content plan; question-shaped headings map straight onto them. Medium

Agentic Readiness · 189 checks

Progress 0 / 189 completed

Hide completed

1 · Discoverability and agent access

0/33

1.1 OAI-SearchBot allowed in robots.txt This is ChatGPT's search and citation crawler, which is not the same thing as GPTBot, its training crawler. Verify: curl https://site.com/robots.txt. High

1.2 A deliberate decision has been made on GPTBot Allowing it means being represented in model training; blocking it means losing long-term parametric knowledge about you. "We haven't decided" is not a decision. Verify: robots.txt plus a written decision note. Medium

1.3 Google-Extended allowed This governs Gemini and AI Overviews. Verify: robots.txt. High

1.4 An explicit position taken on the other agent crawlers ClaudeBot and Claude-SearchBot, PerplexityBot, Amazonbot, Applebot-Extended, Bytespider and CCBot. Verify: robots.txt. Medium

1.5 Agent browsers are not blocked ChatGPT Atlas, Perplexity Comet and Gemini-in-Chrome are user sessions, not bots. Aggressive WAF rules cut them off along with the scrapers. Verify: a live test from an agent browser. High

1.6 WAF and bot management do not block legitimate agent traffic Cloudflare, Akamai, Sucuri or AWS WAF, and in particular the blanket "SI Scrapers" toggle, are set deliberately rather than left at the default. Verify: WAF log review and user-agent testing. High

1.7 Rate limiting does not throttle agents 429 and 403 responses stay below 1% for agent user-agents. Verify: server logs for the last 30 days. Medium

1.8 Newer crawl controls configured deliberately Cloudflare pay-per-crawl and similar bot-charging features. Verify: CDN dashboard. Low

1.9 The page renders without JavaScript Product name, price, stock status and description are present in the HTML source. Most agents run no JavaScript, or only part of it. Verify: curl -s URL | grep price, or a browser with JavaScript disabled. High

1.10 Server response time is under the agent timeout Target TTFB is 600 ms. Verify: PageSpeed or server metrics. Medium

1.11 XML sitemap is current and complete It covers every product page and lastmod values are accurate. Verify: Search Console. Medium

1.12 IndexNow or an equivalent instant-notification method is in place This is what propagates price and stock changes quickly. Verify: Bing Webmaster or the IndexNow API log. Low

1.13 404s, soft 404s and redirect chains are clean The link an agent follows should not be broken. Verify: crawler report. Medium

1.14 No cloaking What the bot is served matches what the user is served. This is content parity, not an optimisation. Verify: fetch comparison using a bot user-agent. High

1.15 Geo-blocking does not cut off agent infrastructure Most agent infrastructure sits in US data centres, so IP blocks aimed elsewhere can silently remove you from your export markets. Verify: an access test through a VPN in the target market. High

1.16 No login wall or cookie wall in front of product information If the agent has to accept something before it can read the price, it will not read the price. Verify: anonymous session test. High

1.17 Agent-facing sitemap published A separate index such as /sitemap-agents.xml collects the product, pricing and policy pages an agent should read first, and robots.txt points at it. Verify: curl https://site.com/sitemap-agents.xml plus the robots.txt line. Medium

1.18 LCP target under 1.2 seconds TTFB alone is not enough: agent browsers (Atlas, Comet) render the page like a real user, and a slow product area times out. Verify: CrUX field data, 75th percentile. High

1.19 HTTPS enforced on every page with no mixed content Agents assign low trust to insecure connections; a single HTTP resource triggers mixed-content warnings. Verify: browser security panel + SSL test. High

1.20 Canonical tags point at the right target on every page Duplicate URLs (parameters, www, trailing slash) must resolve to one canonical; agents read which URL is authoritative from it. Verify: rel=canonical in source. High

1.21 Mobile viewport declared and pages mobile-friendly Agent browsers mostly render with a mobile profile; a page without a viewport tag reads as shrunken desktop. Verify: meta viewport + mobile test. Medium

1.22 nosnippet and max-snippet do not block quotation robots.txt decides whether an agent may enter; these decide what it may show afterwards. nosnippet, or max-snippet:0, means AI Overviews and chat surfaces cannot quote you at all, and max-snippet with a low cap leaves too little context to make the answer. Use -1 for unlimited. Verify: meta robots and the X-Robots-Tag response header. High

1.23 Image previews are not switched off max-image-preview:none and noimageindex keep your product out of image answer cards and shopping surfaces. large is the setting that lets a visual answer include you. Verify: meta robots and X-Robots-Tag. Medium

1.24 No accidental noindex on pages meant to be cited A page outside the index cannot be cited however correct everything else is. Check the header as well as the meta tag, and watch for bot-scoped syntax like "X-Robots-Tag: bingbot: noindex" which is easy to miss. Verify: curl -I plus the page source. High

1.25 A deliberate position on Content-Signal and TDM reservation Cloudflare's Content-Signal lines in robots.txt (search, ai-input, ai-train) and the EU TDM reservation at /.well-known/tdmrep.json are policy declarations, not technical blocks. Publishing them by default, without deciding the scope, quietly gives away visibility. Verify: robots.txt and the well-known file. Low

1.26 One H1 and a heading hierarchy that does not skip levels Zero H1s leaves the topic unstated; several blur which one it is. Agents read the heading tree to work out what answers what. Verify: page source. Medium

1.27 Title and meta description fit without truncation Roughly 25-65 characters for the title and 100-170 for the description. Without a description the engine writes its own summary, and the sentence it picks may not be the one you want. Verify: page source. Medium

1.28 html lang declared and hreflang pairs are bidirectional Without a language declaration the engine has to infer which market you belong to, and one-way hreflang lets language versions cannibalise each other. Verify: the html tag and the link rel=alternate set. Medium

1.29 Cache-Control is declared deliberately With no header, intermediaries guess how long to hold the response and repeat requests pay needless latency. Verify: curl -I. Low

1.30 Contact and privacy pages are linked and reachable A site that leaves unclear who is behind it and how it handles data starts behind on any trust assessment. Verify: footer links returning 200. Medium

1.31 What agents receive matches robots.txt Every agent allowed in robots.txt also gets a 200 from the server. The most common silent failure is the gap between the two: robots.txt is spotless while the WAF returns 403 to the same agent, and a tool that only reads robots.txt calls the site open. Verify: send one request per user-agent and read the FINAL status after redirects; never count a 301 as open. High

1.32 Agent tests include a control request When an agent gets a 403, it is established whether the block is SI-specific or the site simply closes all datacenter traffic. Send the same request with an ordinary browser user-agent: if that is blocked too, the problem is not about agents and the fix is different. Verify: control request plus agent request, same URL, same run. Medium

1.33 Agent access is re-verified on a schedule WAF rules, bot signature lists and agent user-agents change; a one-off check goes stale within months. The access test is repeated regularly and the result is recorded. Verify: a monthly test log and a comparison against the previous run. Medium

2 · Machine readability and structured data

0/29

2.1 Organization schema on the homepage name, url, logo, sameAs, contactPoint and vatID or taxID. Verify: Rich Results Test. High

2.2 Product schema on every product page name, description, image, sku, gtin13 or mpn, and brand. Verify: Rich Results Test. High

2.3 Offer schema complete price, priceCurrency in ISO 4217, availability, priceValidUntil, itemCondition and url. Verify: Rich Results Test. High

2.4 OfferShippingDetails is machine-readable Delivery time, cost and destination region. Verify: schema validator. Medium

2.5 MerchantReturnPolicy is structured Return window, cost and method. Agents treat return risk as a selection criterion. Verify: schema validator. Medium

2.6 AggregateRating and Review reflect real reviews Fabricated reviews carry a permanent trust penalty, not a temporary one. Verify: schema validator plus the review source. Medium

2.7 FAQPage or QAPage on product and category pages Verify: schema validator. Low

2.8 BreadcrumbList and a coherent category hierarchy Verify: schema validator. Low

2.9 Variants are modelled correctly ProductGroup with hasVariant, or isVariantOf. Colour and size confusion is where agents fail most often. Verify: schema validator. High

2.10 Schema matches the visible content exactly A price that differs between the markup and the page reads as a spam signal. Verify: manual comparison across 10 sample products. High

2.11 JSON-LD is the format in use Not microdata, not RDFa. Verify: page source. Medium

2.12 Technical specifications are text, not images Specs belong in a table or list; anything baked into a graphic is invisible to an agent. Verify: page review. High

2.13 Image alt text is descriptive and factual Verify: crawler report. Medium

2.14 /llms.txt is published Company summary, category links and policy links. Verify: curl https://site.com/llms.txt. Low

2.15 The team knows llms.txt is a proposal, not a guarantee It is worth having for multi-engine hygiene, but no promises should be built on it. Verify: team briefing. Low

2.16 Product descriptions are fact-based Dimensions, materials, compatibility and conditions of use, rather than marketing copy. Verify: a 10-product sample review. High

2.17 Comparison and fit information is present Agents eliminate options. A brand that never writes down what it is not suitable for gets penalised for the wrong matches. Verify: content audit. Medium

2.18 PDF and catalogue content has been moved into HTML Verify: site crawl. Low

2.19 Cross-border trade terms machine-readable in schema Duty responsibility (DDP/DDU), IOSS/VAT number, HS code and delivery-region limits are written into Offer/OfferShippingDetails. 5.8 checks that the tax maths is right; this item checks the agent can see it before ordering. Verify: Rich Results Test plus manual review of the product schema. Medium

2.20 Semantic HTML in use; headings, lists and sections are marked up main, article, nav and header let an agent segment the page; a div pile leaves no hierarchy in the content. Verify: source review. High

2.21 A unique meta description on every page 140-160 characters, an honest summary of the page. Agents and search engines read it when classifying quickly. Verify: crawler report. Medium

2.22 Open Graph and social sharing tags defined og:title, og:description, og:image; keeps the content represented correctly on share and preview surfaces. Verify: source + share preview test. Low

2.23 AGENTS.md published and current A machine-readable guide describing how agents should interact with the site; a young convention, no guarantee, but cheap to maintain. Verify: curl https://site.com/AGENTS.md. Low

2.24 skill.md / agent capability definition considered A definition file for the functions the site offers agents; until the ecosystem settles, a deliberate decision is enough. Verify: file presence + decision note. Low

2.25 llms-full.txt full-content version served The expanded companion to llms.txt; lets an agent read the site from a single file. Verify: curl https://site.com/llms-full.txt. Low

2.26 Every JSON-LD block parses A broken block is silently ignored. Nothing on the page shows it and the schema is treated as absent, so the check that matters is validity rather than presence. Verify: the Rich Results Test or a JSON validator on each block. High

2.27 Author is declared as a Person On content claiming expertise, author identity is the most concrete input to a trust assessment. Verify: the author field in the Article schema. Medium

2.28 sameAs links the brand to verifiable profiles Wikipedia, LinkedIn, Crunchbase and the like. This is one of the strongest signals in entity resolution and what keeps you apart from a same-named brand. Verify: sameAs in the Organization schema. Medium

2.29 HowTo schema on step-by-step content Where the page genuinely describes a procedure. Do not force it onto content that is not a procedure. Verify: page source. Low

3 · Product data and feed readiness

0/20

3.1 Required fields are complete id, title up to 150 characters, description up to 5,000 characters in plain text, link, image_link, availability, price as an amount plus ISO 4217 currency, and brand. Verify: feed validator. High

3.2 id is unique and stable It does not change when a campaign runs or stock turns over. Verify: feed diff check. High

3.3 GTIN coverage is at or above 90% 8 to 14 digits, no dashes or spaces. Where there is no GTIN use mpn; where there genuinely is none, set identifier_exists=no. Verify: feed analysis. High

3.4 availability uses the correct enum in_stock, out_of_stock, pre_order or backorder. Verify: feed validator. High

3.5 is_eligible_search and is_eligible_checkout are set deliberately Checkout eligibility requires search eligibility to be true. Verify: feed review. High

3.6 Variants are grouped with item_group_id colour, size and material are populated. Verify: feed analysis. Medium

3.7 google_product_category and product_type are assigned correctly Verify: feed analysis. Medium

3.8 Shipping fields are populated shipping, shipping_weight and delivery time. Verify: feed analysis. Medium

3.9 Enrichment fields are used star_rating, review_count and q_and_a are what separate otherwise identical listings. Verify: feed analysis. Low

3.10 seller_privacy_policy and seller_tos are populated where checkout is active Verify: feed review. High

3.11 Format and encoding are valid UTF-8, delivered as .jsonl.gz, .csv.gz, .tsv.gz or Parquet. Verify: file check. High

3.12 The feed is pushed, not scraped Delivery to an SFTP endpoint rather than leaving it to be crawled. Verify: cron or job log. High

3.13 Freshness targets are met The feed updates within 15 minutes of a price or stock change, with a full snapshot at least daily. Verify: update log. High

3.14 Feed and site prices and stock match exactly A mismatch is a de-ranking cause, not a cosmetic issue. Verify: automated comparison script. High

3.15 Feed error rate stays below 2% Rejection reasons are tracked: missing fields, wrong types, invalid URLs, duplicate ids, credentials in URLs. Verify: validation report. Medium

3.16 Feeds are separated by market Language, currency, price, shipping and eligibility differ per target market. Verify: feed inventory. High

3.17 English and target-language copy has been edited, not machine-translated Verify: sample language review. High

3.18 Prohibited and restricted products have been screened Both target-market law and platform policy. Verify: category review. Medium

3.19 Feed generation is automated It does not depend on someone maintaining a spreadsheet. Verify: process document. Medium

3.20 Feed versioning and rollback are possible A bad feed can be reverted. Verify: version repository. Low

4 · Protocol and integration layer

0/16

4.1 chatgpt.com/merchants application submitted or approved Verify: application record. Medium

4.2 ACP Product Feed Spec compliance is complete This is Part 3 in full. Verify: validation output. High

4.3 A decision has been made on ACP Agentic Checkout Native checkout was pulled back after March 2026, so this is now optional rather than critical. The check is that no oversized investment was made in it. Verify: written decision. Medium

4.4 A UCP manifest is published at /.well-known/ucp Capability declarations are accurate. Verify: curl https://site.com/.well-known/ucp. Medium

4.5 A protocol abstraction layer exists Feed and API are generated from one source and bridged to several protocols, so there is no lock-in to a single one. Verify: architecture document. High

4.6 Payment-protocol developments are being tracked AP2, Visa TAP and Mastercard Agent Pay, with a roadmap obtained from the payment provider. Verify: provider correspondence. Medium

4.7 An MCP server has been evaluated This is how internal data and tools get exposed to agents safely. Verify: technical assessment note. Low

4.8 API-first or headless architecture is in place or on the roadmap Verify: architecture document. Medium

4.9 The commerce platform's agentic support is confirmed in writing Shopify, IdeaSoft, Ticimax, T-Soft, WooCommerce or custom. A sales conversation is not confirmation. Verify: written vendor response. High

4.10 Real-time sync with the ERP or stock system is live Logo, Netsis, SAP or equivalent. Verify: integration log. High

4.11 Webhook and status notification flows work Order status is reported back to the agent. Verify: test transaction. Medium

4.12 A sandbox or test environment exists Nothing is trialled in production. Verify: environment inventory. Medium

4.13 Protocol versions are tracked ACP and UCP spec changes are reviewed quarterly and someone is named as responsible. Verify: tracking record. Medium

4.14 Contextual deep linking supported An agent can send the user to a product or cart URL prepared with the chosen variant, quantity and market rather than to the homepage, and the link works without an existing session. Verify: a variant-parameterised URL producing the correct cart in an anonymous session. Medium

4.15 Structured data ready for function calling Product, price and stock served in a consistent shape agents can consume through function calls. Verify: schema + API response review. Medium

4.16 Decision made on OpenAI Actions / plugin surfaces Whether to expose services to ChatGPT Actions-style surfaces must be a deliberate call; "never looked at it" is not a decision. Verify: decision note + manifest if any. Low

5 · Payments, commercial terms and unit economics

0/11

5.1 You know who the merchant of record is Under ACP it stays with the seller, which means returns, disputes and tax liability are yours. This should be explicit in the contract. Verify: contract review. High

5.2 Commission impact has been modelled Platform commission plus payment processing, expressed as one effective rate and compared against margin. Verify: unit economics sheet. High

5.3 You know which products are profitable in this channel Decided at SKU level. Low-margin lines are not opened to the channel. Verify: margin analysis. High

5.4 Shared payment token or delegated payment flow tested Tested with the payment provider, not assumed. Verify: test transaction record. Medium

5.5 Chargeback liability is settled for agent-originated transactions Verify: bank or PSP correspondence. High

5.6 Local payment methods are supported in the target market Verify: checkout test. Medium

5.7 Multi-currency and FX handling are in place Verify: system check. Medium

5.8 Tax calculation is correct per market VAT, sales tax, customs thresholds, and the IOSS and DDP or DDU decision. Verify: tax adviser sign-off. High

5.9 Price transparency holds end to end The price the agent sees matches the total at checkout. A surprise fee is a cancelled order and a trust penalty. Verify: end-to-end test. High

5.10 Return and cancellation costs are modelled for this channel Verify: financial model. Medium

5.11 Negotiation and dynamic-offer limits are written down Floor price, discount authority and volume tiers for agent-to-agent negotiation are defined as rules, and the algorithm is technically prevented from going below cost. Verify: pricing rules document plus a floor-price test. Low

6 · Operations, inventory and service reliability

0/10

6.1 Real-time stock accuracy is at or above 98% Agents do not tolerate uncertainty. One wrong answer costs trust that does not come back quickly. Verify: physical count against system. High

6.2 Oversell protection is active The same unit is not sold across several channels. Verify: channel reservation logic. High

6.3 Delivery promises are realistic and met Within one day of what was promised. Verify: carrier performance report. High

6.4 The return policy is machine-readable and matches practice Verify: schema plus the policy page. High

6.5 Order status is queryable via API or webhook Verify: test. Medium

6.6 Operations can absorb agent-driven volume Verify: capacity note. Medium

6.7 A kill switch exists for bad price or stock data Automatic halt, tested rather than documented. Verify: procedure plus test. High

6.8 Customer service recognises agent-originated customers And knows how the process differs. Verify: training record. Medium

6.9 Product imagery meets the bar At least one high-resolution white-background image on a permanent URL. Verify: image audit. Medium

6.10 A return address and logistics solution exist in the target market Verify: logistics contract. Medium

7 · Trust, identity and governance

0/10

7.1 An agent authentication approach is defined Know Your Agent thinking: signed requests, allowlists, verified user-agents. Verify: security document. Medium

7.2 Legitimate agents and malicious bots are distinguished at the WAF Verify: WAF rule set. High

7.3 Human approval thresholds are written down Classified by amount, product type and risk level into automatic, approved or forbidden. Verify: policy document. High

7.4 A just-in-time approval mechanism exists Transactions over the limit stop and request approval. Verify: system test. Medium

7.5 Every agent transaction is auditable Who, when and under what authority, retained for at least 12 months. Verify: log sample. High

7.6 A rollback procedure is defined How a mistaken agent order gets cancelled. Verify: procedure document. High

7.7 An incident response plan exists and has been rehearsed Wrong price published, cascading bad orders, prompt injection. Verify: drill record. Medium

7.8 Content and user-generated fields are reviewed for prompt injection and data poisoning Verify: security review. Medium

7.9 Guardrails are written down The list of things an agent may never do, such as creating bulk discounts or sharing customer data. Verify: policy document. High

7.10 Liability is shared explicitly with third-party agents and intermediaries Verify: contract. Medium

8 · SI visibility (GEO and AEO)

0/20

8.1 A prompt set is defined 20 to 50 real purchase-intent prompts in the target market's language, written down and frozen so measurements stay comparable. Verify: prompt list. High

8.2 A baseline measurement exists How many of those prompts mention the brand at all, expressed as a found rate. Verify: measurement report. High

8.3 Position quality is tracked Whether the brand appears early in the answer or at the end of a list. Verify: measurement report. Medium

8.4 Share of model is compared against competitors Verify: measurement report. Medium

8.5 Citation rate is tracked How often the site is given as a source in longer answers. Verify: measurement report. Medium

8.6 Measurement covers several engines ChatGPT, Gemini and AI Mode, Perplexity, Copilot and Claude. Verify: measurement scope. Medium

8.7 Brand entity is unambiguous Consistent naming, a Wikidata or Wikipedia presence, LinkedIn and Crunchbase, and sameAs links that tie them together. Verify: entity audit. Medium

8.8 The brand exists in third-party authority sources Industry lists, comparison sites, reviews and forum or Reddit mentions. Language models lean on these heavily. Verify: source inventory. High

8.9 Review platform profiles are real and current Trustpilot, Google and marketplaces. Verify: profile check. Medium

8.10 Comparison and alternative content is published Agents compare. If you never publish a comparison, someone else's is used. Verify: content inventory. Medium

8.11 Content exists in the target market's language Turkish content does nothing for an agent answering a US buyer. Verify: content language distribution. High

8.12 The limits of the measurement are understood Model answers vary between sessions, so a single measurement proves nothing. Measurements are repeated. Verify: measurement protocol. Medium

8.13 The Matthew effect is accounted for A small brand takes time to earn citations, and short-term targets are set accordingly. Verify: targets document. Low

8.14 Each paragraph reads on its own An answer engine lifts a passage out of its context. A paragraph that only makes sense after the two above it cannot be quoted. Write so that the subject is named rather than referred to as "it". Verify: read any paragraph alone and check it still stands. High

8.15 The opening sentence answers the question directly The lede should state the answer, not promise one. Aim for 8-45 words: too short carries no claim, too long will not fit a quote. Verify: read the first sentence alone and see whether it answers the page's title. High

8.16 Claims carry figures rather than adjectives A verifiable number outranks a general statement. "Much faster" does not distinguish you from the hundreds saying the same thing; "38% faster, measured over 90 days" does. Verify: count the concrete figures on the page. High

8.17 Comparisons are tables and steps are numbered lists Tables and lists are the formats engines parse most reliably and transfer straight into answer cards. Prose that hides a comparison inside a paragraph loses that. Verify: page source. Medium

8.18 Subheadings carry anchor ids Passage-level citation needs an address. A heading without an id cannot be linked to directly, so the whole page is cited instead of the passage that answers the question. Verify: id attributes on h2 and h3. Medium

8.19 Publication and update dates are machine-readable time[datetime], article:published_time or JSON-LD datePublished. Undated content is dropped on questions where freshness matters, however good it is. Verify: page source. Medium

8.20 Claims are tied to outside sources Linking out to what a claim rests on reads as a credibility signal; a page that cites nothing reads as weakly verifiable. Verify: count the outbound domains in the content area. Low

9 · Measurement and attribution

0/10

9.1 A custom GA4 channel group isolates AI sources Built with session source matches regex, and ordered above Referral so it wins. Verify: GA4 admin screen. High

9.2 The regex covers the main sources chatgpt\.com, chat\.openai\.com, openai\.com, perplexity\.ai, gemini\.google\.com, claude\.ai, copilot\.microsoft\.com, you\.com and meta\.ai. Verify: rule review. High

9.3 Dark traffic is disclosed in the reporting A large share of SI traffic carries no referrer and lands in Direct. The numbers are presented as a floor, not a total. Verify: report template. High

9.4 A second source exists Server-side tagging or server log analysis. Verify: installation check. Medium

9.5 UTM taxonomy is standard and consistently applied Verify: taxonomy document. Medium

9.6 The funnel from AI session to sale is traceable Verify: GA4 report. High

9.7 Visibility, traffic and revenue meet in one dashboard Verify: dashboard. Medium

9.8 There is a monthly reporting rhythm with a named owner Verify: calendar. Medium

9.9 The baseline date is recorded Without it there is no before-and-after. Verify: baseline report. High

9.10 Third-party conversion claims are not repeated without local proof Figures like "SI traffic converts X times better" are validated against your own data first. Verify: report note. Medium

10 · SI advertising (GEA) readiness

0/10

10.1 Geographic eligibility is clear ChatGPT Ads does not target Turkiye. The open markets are the US, the UK, Canada, Australia, New Zealand, Japan and South Korea. Verify: campaign settings. High

10.2 The legal entity and market position for account opening is settled Verify: legal confirmation. High

10.3 Vertical restrictions have been checked Health, finance and legal are among the restricted categories. Verify: policy reading. High

10.4 A context hints library is prepared Not keywords: descriptions of the situations in which the product is useful, 5 to 15 variants per ad group. Verify: hints document. Medium

10.5 Creative specifications are met Short headline, one-fact description, square image and favicon. Verify: creative set. Medium

10.6 Landing pages match the hint themes Relevance score feeds directly into cost, which is why GEO work lowers the ad bill. Verify: page mapping table. High

10.7 Pixel and Conversions API are installed Verify: measurement test. High

10.8 The Google side is ready Feed and assets prepared for PMax, AI Max and AI Overviews placements. This side does work in Turkiye. Verify: campaign structure. Medium

10.9 Test budget and learning period are planned No decisions are made on early numbers. Verify: media plan. Medium

10.10 A manual A/B protocol is defined The platform may not provide a built-in test primitive. Verify: test plan. Low

11 · Legal, compliance and data

0/10

11.1 KVKK compliance covers agent-originated transaction data Data minimisation, disclosure and explicit consent. Verify: legal opinion. High

11.2 GDPR and the data transfer framework are complete Applies if you sell into the EU. Verify: legal opinion. High

11.3 The 1 August 2026 Commercial Advertising Regulation is met Transparency obligations for SI-generated content and digital characters. Verify: creative review. High

11.4 EU AI Act transparency obligations have been assessed Applies if you export to the EU. Verify: compliance note. Medium

11.5 Target-market consumer law is covered Right of withdrawal, price transparency and auto-renewal rules. Verify: legal check. High

11.6 The moment a contract is formed through an agent is defined Along with how the consent record is stored. Verify: procedure. High

11.7 Product compliance is documented for the target market Product safety, labelling and certification such as CE, UKCA or FDA. Verify: compliance file. High

11.8 Customs, origin and HS codes are correct Verify: export file. Medium

11.9 Intellectual property is clear in the target market Trademark and image usage. Verify: trademark check. Medium

11.10 Contract and log retention periods meet the regulation Verify: retention policy. Medium

12 · Organisation, capability and sourcing

0/10

12.1 One process owner is named Part-time is fine. Anonymous is not. Verify: role description. High

12.2 Decision authority levels are written down SI recommends, human approves, or fully automatic. Verify: process document. High

12.3 The delivery resource is identified In-house, agency or vendor. Verify: resourcing plan. Medium

12.4 Vendor contracts are outcome-based Not "X hours of consultancy" but verifiable milestones: feed validated, measurement installed, campaign live. Verify: contract. High

12.5 The team has been trained At least one person can read a feed and a schema block. Verify: training record. Medium

12.6 A monthly review is in the calendar Verify: calendar. Medium

12.7 Budget is allocated and the return expectation is realistic Verify: budget. Medium

12.8 The delivery sequence is understood Not training then application, but a small real task, then training, then application again. Classroom training alone has weak effect; application with coaching and feedback is what works. Verify: programme design. Medium

12.9 Pre-launch ARI gate cleared At least 80% of product pages are machine-readable without running JavaScript, and the API behind "Buy" answers in under 2 seconds. Agent integration does not go live until both thresholds hold. Verify: JS-disabled sample crawl plus API response-time measurement. High

12.10 ARI score re-measured on a cycle The score is not a one-off certificate: it decays as protocols and platforms move. Measurement repeats quarterly and is compared with the previous score. Verify: quarterly score record. Medium

[Full page with the live deep scan](https://www.webtures.com/si-agent-readiness/)
