- Through September 2026, BeCited graded sites on 19 AI-readiness checks, weighted into a letter grade. Today’s paid audits report only the hard blockers.
- The single biggest technical mistake we see is allowing training bots (GPTBot, Google-Extended) while blocking retrieval bots (OAI-SearchBot, PerplexityBot, Claude-SearchBot). 71% of sites have this misconfiguration.
- Schema.org JSON-LD (12 points) and entity readiness (8 points) were the highest-weighted single checks. Both are usually fixable in a week.
- Core Web Vitals (LCP, INP, CLS) counted for 6 points.
Site readiness is the technical layer of GEO. It is what determines whether AI crawlers can reach your content, parse it cleanly, quote it accurately, and understand who you are as an entity.
Through September 2026, BeCited ran these 19 checks before any questions were asked and turned them into a letter grade. Today’s paid audits report only the hard blockers: a robots.txt that shuts AI out, a noindex tag, content that only appears with JavaScript, or a site AI can’t reach. The free check on our homepage still runs 9 of these signals. The checks fall into seven tiers, mapped to the three pillars of GEO (retrievability, citability, recognizability).
The full list with weights
In that grade, default weights summed to 116 across the 19 checks. Some checks were weighted by business type: local services skipped agentic readiness, and software companies counted it double.
| Check | Tier | Weight |
|---|---|---|
| robots.txt | Crawlability | 8 |
| llms.txt / llms-full.txt | Crawlability | 2 |
| sitemap.xml | Crawlability | 5 |
| Information architecture | Crawlability | 6 |
| Bot identity verification | Crawlability | 3 |
| JSON-LD schema.org | Structured metadata | 12 |
| OpenGraph & meta tags | Structured metadata | 8 |
| Heading structure | Structured metadata | 5 |
| Quotable claims | Content extraction | 4 |
| Semantic HTML | Content extraction | 7 |
| FAQ content format | Content extraction | 3 |
| E-E-A-T signals | Content extraction | 5 |
| Content freshness | Content quality | 8 |
| Quotability score | Content quality | 8 |
| Information gain | Content quality | 7 |
| Rendering completeness | Rendering | 6 |
| Entity readiness | Entity & agent | 8 |
| Agentic readiness | Entity & agent | 5 |
| Core Web Vitals | Performance | 6 |
| Total | 116 |
Tier 1: Crawlability & discovery
If AI crawlers cannot reach your content, nothing else matters. This tier is the foundation of retrievability.
1. robots.txt — AI crawler classification
8 ptsBeCited classifies AI bots as training (GPTBot, Google-Extended, ClaudeBot, CCBot) versus retrieval (OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-SearchBot). Blocking training bots is a defensible policy choice. Blocking retrieval bots makes you invisible to AI search.
The most common anti-pattern: brands block GPTBot to "protect" content but accidentally block OAI-SearchBot too. Per a BuzzStream analysis, 71% of sites have this misconfiguration.
71%
Of sites that block training bots also block retrieval bots like OAI-SearchBot and PerplexityBot, which makes them invisible to AI search.
BuzzStream, 2025
2. llms.txt & llms-full.txt
2 ptsAn emerging standard for handing LLMs a curated index of your most quotable pages. No major answer engine reads it for search today. Its value is in agents and coding tools that read it directly, so the grade weighted it low, as cheap insurance.
3. sitemap.xml
5 ptsThe classic sitemap is still load-bearing for AI crawlers. We check it exists, that it is reachable from robots.txt, and that lastmod values are filled in, since those dates feed the freshness check below.
4. Information architecture
6 ptsThe check looked for topic clusters (pages grouped under shared URL paths), readable URL slugs that match the page’s subject instead of opaque IDs, and a sensible priority order in the sitemap. AI retrieval rewards sites with clear hierarchy; orphan pages and flat URL structures lose ranking against well-clustered competitors.
5. Bot identity verification
3 ptsA forward-looking check based on the IETF Web Bot Auth draft (Cloudflare-led, 2025), which lets a site verify which AI bot is at the door. It asked two things: can the CDN or firewall in front of your site verify signed bot requests, and which AI companies publish their signing keys. The standard is still early, so the check warned but never failed a site.
Tier 2: Structured metadata
Once a crawler has your page, structured metadata tells it what the page is about. In our earlier site grade, this was the highest-weighted tier.
6. JSON-LD schema.org
12 ptsThe biggest single item in our earlier site grade. Independent studies put the citation uplift from proper structured data at 1.8–3.2x. The check looked for type-appropriate schema (LocalBusiness or Service for local; SoftwareApplication for SaaS; Product for consumer goods) plus Organization, FAQPage, AggregateRating, and Review where relevant.
1.8–3.2x
Citation uplift for pages with proper JSON-LD schema.org markup vs. pages without.
Princeton GEO & LSEO
7. OpenGraph & meta tags
8 ptsTitle, description, og:title, og:description, og:image. AI engines often quote the meta description word for word when the page is cited.
8. Heading structure
5 ptsOne H1 per page, descriptive H2s, and a consistent hierarchy. Headings phrased as questions tend to earn 40% more citations because they map directly to user prompts.
Tier 3: Content extraction signals
This tier is research-backed against the Princeton GEO paper and LSEO content extraction studies. It covers whether AI can quote your content.
9. Quotable claims
4 ptsSelf-contained 50–150-word chunks with answer-first structure. AI engines pull blocks, not paragraphs. A page full of conversational prose with no extractable claims will lose to a competitor with one well-formed answer block.
10. Semantic HTML
7 ptsUse main, article, section, lists and tables instead of div soup. Engines parse semantic elements faster and trust them more. We count semantic tags vs unstructured div containers and flag the imbalance.
11. FAQ content format
3 ptsFAQPage schema, native HTML <details> elements, or Q&A-style headings. FAQs are over-represented in AI citations because their structure aligns with the prompt-and-answer pattern of generative search.
12. E-E-A-T signals
5 ptsPerson/author schema, byline markup, credential language ("certified", "licensed", "N years of experience"). Per AI Overview research, 96% of cited pages have strong E-E-A-T signals. A position-6 page with E-E-A-T markup beats a position-1 page without.
96%
Of pages cited in Google AI Overviews carry strong E-E-A-T signals: author markup, credential language, Person schema.
AI Overview research
Tier 4: Content quality
Three checks on the content itself, beyond formatting.
13. Content freshness audit
8 ptsLast-Modified HTTP headers, JSON-LD dateModified, and sitemap lastmod. Perplexity weights content under 30 days old at roughly 3.2x. Half of all AI citations come from content less than 11 months old.
3.2x
Citation weight Perplexity applies to content updated within the last 30 days vs. older content.
Digital Bloom & AirOps
14. Quotability score
8 ptsA composite of paragraph length distribution, answer-first pattern detection, statistic density, and self-contained chunk count. Engines reward content that is easy to lift, and this score measured how easy yours is.
15. Information gain
7 ptsThe Princeton GEO paper identifies originality as the primary differentiator between content that gets cited and content that gets filtered out as paraphrase. We measure first-person research markers ("in our research", "we tested"), original data tables, and case-study evidence, and we discount paraphrase markers ("according to", "studies show") that signal derivative reporting. Pages with high information gain dominate AI citations even when their authority signals are weaker than competitors.
Tier 5: Rendering completeness
AI crawlers (GPTBot, ClaudeBot, PerplexityBot) typically don't run JavaScript, so content that JavaScript adds after the page loads is invisible to them, even if the static HTML looks healthy.
16. Rendering completeness
6 ptsThe check sorted every script tag (render-blocking, deferred, async, inline, JSON-LD), looked for empty app shells (<div id="root"></div>), confirmed that critical content (an H1 and at least 300 body words) is in the initial HTML, and measured image alt-text coverage. An optional second fetch rendered the page with JavaScript and compared word counts, to catch pages whose static HTML looks fine while the visible content is added by JavaScript.
Tier 6: Entity & agent signals
This tier covers whether engines understand who you are at the entity level and, for SaaS, whether AI agents can use your product.
17. Entity readiness
8 ptsWikipedia presence, Wikidata entry, Organization schema with sameAs links to authoritative profiles, and consistent brand naming across the web. AI engines are reluctant to recommend entities they cannot disambiguate.
18. Agentic readiness (SaaS only)
5 ptsAGENTS.md, OpenAPI spec, public API documentation, and MCP manifest. As Anthropic's Model Context Protocol, Google's UCP, and Visa's Agentic Ready standards mature, this tells AI engines that agents can call your product directly.
Tier 7: Performance (Core Web Vitals)
Performance applies to every site, but the grade weighted it more for web apps than for brick-and-mortar marketing pages.
19. Core Web Vitals
6 ptsBeCited's earlier site check measured the three Core Web Vitals with Google PageSpeed Insights: Largest Contentful Paint (LCP), Interaction to Next Paint (INP), and Cumulative Layout Shift (CLS). Pass thresholds (per Google web.dev) are LCP under 2.5s, INP under 200ms, and CLS under 0.1. Warn thresholds are 4s, 500ms, and 0.25; above those, fail.
How the grade was calculated
The grade was a weighted percentage. Each check returned pass, warn or fail. Pass earned full weight, warn earned half, fail earned zero. Points earned divided by points available gave a 0–100 score, which mapped to A (85+), B (70–84), C (55–69), D (40–54), or F (under 40).
A brand can show up well in AI answers and still have a robots.txt that blocks retrieval bots. We still flag that as the top action item.
Where most brands fail first
Across audits we have run, three patterns recur:
- Robots.txt over-blocks. Either the site blocks all AI bots indiscriminately, or (more often) it blocks the wrong subset.
- Schema is missing or wrong-typed. Many sites have JSON-LD for WebSite or BreadcrumbList but nothing for the entity that matters: LocalBusiness, SoftwareApplication, or Product. AI engines cannot ground claims to a missing entity.
- Quotability is low. Long meandering paragraphs, no answer-first structure, no statistics. The competitor with shorter, denser blocks gets quoted; the long-form site gets ignored even when its content is better.
Start with the hard blockers: a robots.txt that shuts out retrieval bots, a noindex tag, or content that only appears with JavaScript. If AI can’t reach or read the page, nothing else on this list matters.
Frequently asked questions
What is site readiness in a GEO audit?
Site readiness is the technical layer of GEO. It is the set of signals that determine whether AI crawlers can reach your content, parse it cleanly, quote it accurately, and understand who you are at the entity level. Through September 2026, BeCited graded audited sites on 19 checks across seven tiers, weighted into a letter grade. Today’s paid audits report only the hard blockers: a robots.txt that shuts AI out, a noindex tag, content that only appears with JavaScript, or a site AI can’t reach.
What is the most common technical mistake brands make?
Allowing training bots like GPTBot and Google-Extended while accidentally blocking retrieval bots like OAI-SearchBot, ChatGPT-User, PerplexityBot, and Claude-SearchBot. A BuzzStream analysis found 71% of sites have this misconfiguration. The retrieval-bot block makes the site invisible to AI search even when the brand is well known.
Why did JSON-LD schema carry the highest weight?
In BeCited's earlier site grade (through September 2026), JSON-LD schema.org markup was weighted at 12 points, the largest single check, because independent studies put the citation uplift from proper structured data at 1.8 to 3.2x. AI engines use schema to ground claims to specific entities. Without LocalBusiness, SoftwareApplication, or Product schema, the engine has no entity to anchor the claim to and is far less likely to recommend the brand by name.
What is the difference between training and retrieval AI bots?
Training bots crawl content for model training. The major ones are GPTBot, Google-Extended, ClaudeBot, and CCBot. Retrieval bots fetch live content to answer a user query in real time. The major ones are OAI-SearchBot, ChatGPT-User, PerplexityBot, and Claude-SearchBot. Blocking training bots is a defensible policy choice. Blocking retrieval bots makes you invisible to AI search at the moment of the query.
Are Core Web Vitals part of AI readiness?
Yes. In BeCited's earlier site grade, Core Web Vitals were check 19, weighted at 6 points. The check measured Largest Contentful Paint (LCP), Interaction to Next Paint (INP), and Cumulative Layout Shift (CLS) with Google PageSpeed Insights. Pages under 2.5 seconds, 200 milliseconds, and 0.1 respectively earned full credit. Performance matters for AI readiness because retrieval bots and human users land on the same page, and a page that times out is invisible to both.
The article lists the checks. The audit tells you what’s blocking you.
See what ChatGPT and Google’s AI tell buyers in your city about your business. The $20 Snapshot asks about 25 of their questions.
See where you stand: $20 Free check See $200 Full AuditSources cited. Robots.txt misconfiguration rate (71%) is from BuzzStream's bot-traffic analysis. JSON-LD citation uplift (1.8–3.2x) aggregates findings from the Princeton GEO paper and LSEO. Heading-question citation uplift (40%) is from internal BeCited audit data corroborated against the Princeton paper. The 96% E-E-A-T statistic is from AI Overview research. Perplexity freshness weighting (3.2x for content under 30 days) and the half-of-citations-under-11-months figure come from Digital Bloom and AirOps.