Readiness checks

On this page

This page lists every check Site readiness runs, how each one is judged, and why it is in the rubric. Where a search engine or AI company documents the behaviour a check relies on, the source is linked. Where a check rests on a published study, the study is linked. Where it rests on our own reasoning, the reasoning is given.

A check stays in the rubric only if it names a documented or clearly reasoned effect on one of two things. Either it affects whether an AI assistant can reach, read, understand or quote your pages, or it affects how a search engine crawls, indexes or shows them. Purely accessibility or general good-practice checks are not scored. Each finding in Site readiness shows the same reasoning under Why it matters.

How the scores work

Each scan produces two scores from the same checks:

  • Site readiness asks whether AI assistants can find, read, trust and quote your site.
  • SEO Health weighs the same evidence the way search engines do.

The checks fall into five pillars, and each pillar carries a fixed share of each score.

PillarSite readinessSEO Health
Access: can crawlers reach the site2520
Structure: can machines read what each page is2030
Answerability: can a passage be quoted as an answer255
Authority: can the source be trusted2015
Delivery: does the page arrive quickly and whole1030

Within a pillar, checks share its points by their weight, listed with each check below. A check that does not apply to your site, or could not be measured on a scan, drops out, and the others in its pillar share the points.

Each check ends in one of these results:

  • Pass: full points.
  • Warning: part of the points.
  • Fail: no points.
  • Not measured: left out of the score, with the reason shown.

Notes are checks with no weight. They report what was found and change no score.

The free site check and full scans

The free site check reads up to six pages. It does not grade your copy with AI, so it leaves Answerability out of both scores and shows the range a full scan would put the score in. Because of that, a free score can sit above the score a later full scan gives.

These checks run only on full scans in the app:

  • near-duplicate titles and descriptions
  • sitemap hygiene
  • schema that fits each page type
  • reachability from your market
  • the AI-graded Answerability checks
  • summary near the top
  • lists and tables
  • answers to your customers' questions

Access

AI crawler directives in robots.txt

Checks that your robots.txt lets the crawlers AI assistants use to read and cite pages reach your homepage and the pages you leave open to every crawler. These are OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot, Bingbot and Grok-User.

  • Fails when one of them is blocked from the homepage, or from a quarter or more of the pages checked.
  • Warns when one is blocked from some pages. The finding lists them.
  • Passes when none is blocked. Blocking only training crawlers, such as GPTBot, ClaudeBot or Google-Extended, also passes: it is your choice and does not stop assistants quoting you. A site with no robots.txt passes too, because every crawler may read it.
  • Pages that robots.txt closes to every crawler are left out: that is a choice about all crawlers, not an AI block.
  • The suggested fix unblocks only the blocked agents and changes nothing else in the file. If your hosting or security provider generates robots.txt, the fix points to its dashboard instead.
  • Weight: 10 in Site readiness, 5 in SEO Health.
  • Why: OpenAI, Anthropic and Perplexity each document the user agents their search and answer features use and that they follow robots.txt, so a disallow keeps a page out of what they can quote. Source: OpenAI crawlers, Anthropic crawlers, Perplexity crawlers.

Robots meta and snippet directives

Checks for noindex, nosnippet and max-snippet:0 in each page's robots meta tag or its X-Robots-Tag header. Legal pages and account, sign-in and basket pages are skipped, because they are usually hidden on purpose.

  • Fails when the homepage carries one, or a quarter or more of the pages checked do.
  • Warns when some pages do. The finding says whether each comes from the markup or a header.
  • noai and noimageai pass with a note: they ask not to be used for training, and no major engine reads them when answering.
  • Weight: 3 in Site readiness, 4 in SEO Health.
  • Why: Google documents that noindex keeps a page out of results, and that nosnippet and max-snippet:0 stop text from it being shown, including in its AI features. Source: Robots meta tag.

Homepage status and redirects

Checks that the homepage answers first time.

  • Fails when the homepage answers with a server error, a 404 or a 410. It also fails when the homepage answers 200 but its title or main heading reads like an error page and almost nothing else is on it (a soft 404).
  • Warns when reaching the homepage takes two redirects, and more strongly at three or more.
  • If your site refused our scanner but the page could be read another way, the finding says so and does not count it against you. Refusing our scanner does not mean AI assistants are blocked.
  • Weight: 2 in Site readiness, 4 in SEO Health.
  • Why: Google documents that a page answering with an error is not indexed, and that a page answering 200 with error content is treated as a soft 404. Source: HTTP status codes and network errors.

Checks each page for a gate that hides its content. A gate is judged by its effect: a cookie or region notice counts as a wall only when its wording is in the visible text and fewer than 1,200 characters of readable content sit beside it.

  • Fails when the homepage, or a quarter or more of the pages checked, shows a cookie wall or region gate.
  • Warns when some pages do.
  • Passes when a cookie banner sits over readable content.
  • A paywall never fails. The finding says that assistants can see the page but cannot quote what is behind the wall, and whether the paywall is declared the way search engines expect.
  • Weight: 2 in Site readiness, 3 in SEO Health.
  • Why: content a crawler cannot see without an interaction is not indexed, and Google documents how a paywall should be marked so it is not mistaken for cloaking. Source: Paywalled content.

Reachable from the market you sell to

Full scans only. Reads your homepage as our crawler through a connection in your market's country, and compares its visible text with what we read directly.

  • Fails when the market gets no readable content or an error.
  • Warns when it gets less than 60% of the content, or lands on a different hostname. If the difference is deliberate, for licensing, regulation or localisation, nothing needs changing.
  • Weight: 4 in Site readiness, 2 in SEO Health.
  • Why: an assistant answering for a market reads the page that market is served. If that page is empty or much thinner than the one we see, there is less of yours to quote there.

Domain serves a real site

Checks that the domain serves a working site, not a parked, for-sale, expired, suspended, challenge or placeholder page.

  • Fails when the domain is parked, listed for sale, expired or suspended, or serves only a challenge page.
  • Warns when the homepage shows signs of an unfinished site, such as coming-soon or under-construction wording on an otherwise thin page.
  • Weight: 3 in Site readiness, 3 in SEO Health.
  • Why: a parked, expired, suspended or placeholder domain has nothing to index or quote, so every other finding waits on it.

HTTPS works with a valid certificate

Checks that the site is served over HTTPS with a valid certificate, and that plain HTTP redirects to it.

  • Fails when there is no HTTPS, or the certificate is expired or issued for another name, or HTTP does not redirect to HTTPS.
  • Warns when the redirect to HTTPS takes several hops, or the certificate expires within 14 days.
  • Weight: 2 in Site readiness, 3 in SEO Health.
  • Why: Google uses HTTPS as a page experience signal, and browsers and crawlers that check certificates refuse a connection whose certificate is expired or for another name. Source: Page experience.

Redirect chains

Checks internal links that take more than one redirect to reach their page.

  • Warns when any do, more strongly with 10 or more chains or a chain of four or more hops. This check never fails.
  • Weight: 1 in Site readiness, 2 in SEO Health.
  • Why: Google documents following redirects and recommends pointing to the final address; each extra hop is another request before the page is read. Source: Redirects and Google Search.

Content-Signal declarations (note)

Reports whether robots.txt carries a Content-Signal line stating whether your content may be used for search, AI answers and AI training. The suggested line is one you add to your existing User-agent: * group, with the values you choose.

  • Why it changes no score: Content-Signal is a newer, voluntary line that few services read yet. Source: contentsignals.org.

llms.txt (note)

Reports whether /llms.txt exists, is well formed and its links load.

  • Why it changes no score: llms.txt is a proposed file listing a site's most useful pages for language models, and few assistants read it yet. Source: llmstxt.org.

Structure

Schema.org coverage

Checks how many pages carry structured data saying what they are, such as Organization, LocalBusiness, Product, Service, Article, FAQPage or Event. Blocks that only name the site or page, or list breadcrumbs, do not count.

  • Passes when 80% or more of the pages do.
  • Warns when some do.
  • Fails when none does.
  • Weight: 2 in Site readiness, 5 in SEO Health.
  • Why: Google documents that structured data helps it understand what a page is about and makes pages eligible for rich results. Source: Intro to structured data.

Schema validity and placeholder values

Checks that every structured data block parses and holds your real details, not setup text such as {{shop_name}} left from a template.

  • Weight: 4 in Site readiness, 4 in SEO Health.
  • Why: Google documents that structured data must be valid and match the visible page; a block that does not parse is ignored whole. Source: Structured data policies.

Organization schema completeness

Checks the homepage for an Organization or LocalBusiness block with a name, URL and description. The finding says what is missing and links Google's guide; Cleotic does not write the markup for you.

  • Fails when there is no organisation block.
  • Warns when it is missing one of the three.
  • Weight: 3 in Site readiness, 3 in SEO Health.
  • Why: Google documents Organization markup as the way to tell it a business's name, logo, address and other details. Source: Organization structured data.

H1 presence and uniqueness

Checks that each page has a main heading with text.

  • Warns when pages have none, more strongly when over half do.
  • Several H1s on a page pass with a note.
  • Weight: 3 in Site readiness, 4 in SEO Health.
  • Why: the main heading is the clearest statement of what a page is about, for readers, screen readers and search engines.

Canonical correctness

Checks each page's canonical link.

  • Fails when a canonical points to a page the scan read that redirects, answers with an error or asks not to be indexed.
  • Warns when a page has no canonical, or names an address on another domain. That can be deliberate, for syndication or domain consolidation.
  • Weight: 2 in Site readiness, 5 in SEO Health.
  • Why: Google documents rel=canonical as how a site says which of several similar addresses is the one to index. A target that redirects, errors or is noindexed sends a contradictory signal. Source: Consolidate duplicate URLs.

Entity consistency across markup and copy

Checks that your business name, phone number and address read the same in structured data and in the visible copy.

  • Warns for each of the three that disagrees.
  • Weight: 3 in Site readiness, 3 in SEO Health.
  • Why: when the markup and the copy disagree, anything reading both has to pick one, and it may pick the wrong one.

Sitemap completeness

Checks for a valid XML sitemap that lists the pages the scan found, with dates.

  • Warns when there is no valid sitemap. That is fine for a small site whose pages all link to each other.
  • Warns when more than three crawled pages are missing from the sitemap, or most entries have no lastmod date.
  • Weight: 3 in Site readiness, 5 in SEO Health.
  • Why: Google documents sitemaps as how it finds pages, especially new or deep ones on a large site that internal links do not reach quickly. Source: Sitemaps overview.

Locale declarations (hreflang)

Checks hreflang sets across the pages the scan read.

  • Fails when a set is missing return links, judged only where both pages were read, or uses invalid language or region codes.
  • Warns when an alternate redirects, errors or asks not to be indexed.
  • A missing x-default is a note.
  • Weight: 0 in Site readiness, 3 in SEO Health.
  • Why: Google documents hreflang for telling it about language and regional versions of a page, and ignores a set without return links or with invalid codes. Source: Localized versions of your pages.

Title and meta description

Checks each page's title and meta description.

  • Fails when a page has no title, or every page has the same one.
  • Warns when pages have no description.
  • Title and description lengths are a note.
  • Weight: 2 in Site readiness, 6 in SEO Health.
  • Why: Google documents that it uses the title element for the title link in results and often the meta description for the snippet. Source: Title links.

Document language declaration

Checks that each page states its language in the lang attribute.

  • Fails when no page does.
  • Warns when some pages do not, or a value is not a recognised language code.
  • Weight: 0 in Site readiness, 1 in SEO Health.
  • Why: some search engines, Bing among them, read the declared language to decide which language and region a page is shown for. Google works out the language from the visible text, so this weighs only in SEO Health.

Checks that links are real <a href> links, not scripts.

  • Fails when more than half the links on a page run only in JavaScript.
  • Warns when some do.
  • Generic link text such as "click here" is a note.
  • Weight: 1 in Site readiness, 2 in SEO Health.
  • Why: Google documents that it follows only links in an a element with an href, and uses link text to understand the page linked to. Source: Crawlable links.

Internal linking

Checks how your pages link to each other.

  • Fails when the homepage links to fewer than three of your other pages.
  • Warns when other pages do.
  • Pages nothing links to count only when the scan reached every address in your sitemap; otherwise they are a note. The free site check does not report them.
  • Weight: 0 in Site readiness, 2 in SEO Health.
  • Why: Google documents that it finds most pages through links from pages it already knows; a page nothing links to is found late or not at all. Source: Crawlable links.

Character set declared

Checks that each page names its character set near the top.

  • Warns when a page with non-ASCII text, such as accented letters, currency signs or other scripts, does not.
  • A plain-ASCII page without one is a note.
  • Weight: 1 in Site readiness, 1 in SEO Health.
  • Why: a page that does not declare its character set leaves the reader to guess it, and a wrong guess garbles every non-ASCII character in what is indexed or quoted. Source: W3C: declaring character encodings.

Favicon

Checks for an icon declared by the homepage or served at /favicon.ico.

  • Warns when there is none.
  • Weight: 0 in Site readiness, 1 in SEO Health.
  • Why: Google shows a site's favicon beside its results. Source: Favicon in search results.

Checks links to your own pages that answer 404, 410 or a server error. Links that answered 401, 403 or 429, or timed out, are listed as not checked and cost nothing. Links into account, sign-in and app pages are skipped.

  • Fails at 10 or more broken links.
  • Warns at fewer.
  • Weight: 2 in Site readiness, 3 in SEO Health.
  • Why: Google documents that a 404 or 410 page is dropped from the index, so a link to one sends readers and crawlers nowhere. Source: HTTP status codes and network errors.

Near-duplicate page titles

Full scans only. Checks for pages whose titles are the same once the brand name, case and punctuation are set aside, or share nine tenths of their words. Paginated pages, listing pages and pages whose canonical points elsewhere are left out.

  • Fails when 30% or more of pages are affected.
  • Warns below that.
  • Weight: 0 in Site readiness, 2 in SEO Health.
  • Why: Google documents that titles should be distinct for each page, because a title is how a reader tells results apart. Source: Title links.

Near-duplicate meta descriptions

Full scans only. The same test as near-duplicate titles, applied to meta descriptions.

  • Weight: 0 in Site readiness, 1 in SEO Health.
  • Why: Google documents that descriptions should be distinct for each page, because the snippet is how a reader tells results apart. Source: Snippets.

Sitemap hygiene

Full scans only. Checks the addresses your sitemap lists. An address counts against you when it answers 404, 410 or a server error, redirects, or asks not to be indexed. Addresses that answered 401, 403 or 429, or timed out, are listed as not checked.

  • Fails when 20% or more of listed addresses, or 10 or more, are affected.
  • Warns at fewer.
  • Weight: 0 in Site readiness, 2 in SEO Health.
  • Why: Google documents that a sitemap should list the canonical addresses you want indexed. Addresses that error, redirect or ask not to be indexed waste the crawl. Source: Sitemaps overview.

Schema that fits each page type

Full scans only. Checks that pages carry the structured data type that matches what they are, but only where the page shows it:

  • Product on a product page that shows a price, an offer or a buy button.
  • Article on an article address that shows a date or an author.
  • LocalBusiness on a location page with a visible address.
  • BreadcrumbList on pages below the homepage, at half weight.

Results:

  • Warns when 40% or more of the expected types are present.
  • Fails below that.
  • Declaring other types, such as FAQPage or Service, is not penalised.
  • The finding names the type, the reason it is expected, and links Google's guide. Cleotic writes only the BreadcrumbList markup for you.
  • Weight: 0 in Site readiness, 2 in SEO Health.
  • Why: Google documents Product, Article, LocalBusiness and BreadcrumbList markup and the details each makes eligible to show in results. Source: Structured data for ecommerce.

Open Graph and Twitter card (note)

Reports whether the homepage has a complete share card: a title, a summary and a picture that loads.

  • Why it changes no score: these tags decide how a page looks when it is shared, and do not affect search or AI answers.

Answerability

The first three checks are graded by AI, on full scans with plans that include AI grading. Each passage is scored from 1 to 5. The grader also judges whether a passage answers a question. Narrative, testimonial and brand copy is left out of direct answers and fact density rather than marked down.

Graded checks share these bands:

  • Passes when the average is 4 or more.
  • Warns from 2.5.
  • Fails below 2.5.

Each graded finding lists the passages with their score and the grader's reason.

Direct-answer density

Grades whether each section that answers a question gives the answer in its first sentence. When no passage answers a question, the check is not measured.

  • Weight: 4 in Site readiness, 2 in SEO Health.
  • Why: answers are easier to quote when they come first. An assistant lifting a passage takes its opening, and a passage that builds up to its point leaves the answer out.

Chunk self-sufficiency

Grades whether each section names its subject at the start, so it still makes sense when quoted on its own. Pronouns after that are fine.

  • Weight: 3 in Site readiness, 1 in SEO Health.
  • Why: assistants quote passages, not pages. A passage that names its subject still makes sense on its own; one that opens with "it" or "this" does not.

Fact density

Grades how many specific facts, such as prices, times, places and numbers, the answering passages state in place of general claims.

  • Weight: 3 in Site readiness, 1 in SEO Health.
  • Why: the GEO study (Aggarwal et al., KDD 2024) found that concrete figures and statistics raised visibility in generative-engine answers. Source: GEO: Generative Engine Optimization.

Lists and tables machines can lift

Full scans only. Checks that lists are real lists and tables have a header row, rather than grids built from layout elements.

  • Warns or fails depending on how many pages are affected.
  • Weight: 2 in Site readiness, 1 in SEO Health.
  • Why: lists and tables written as real HTML keep each item and figure separate when a page is converted to text. Grids built from layout elements run together.

Answers to the questions your customers ask

Full scans only, and only when you track questions. Checks your tracked questions about what you offer against your headings, FAQ entries and the paragraph that opens each section. The questions cover things like cost, duration, eligibility, process, hours and areas. Comparison and recommendation questions, such as "best", "vs" or "alternatives to", are left out and listed.

  • Passes when half or more are answered.
  • Warns from a fifth.
  • Fails below that.
  • The finding lists the questions no page answers.
  • Weight: 3 in Site readiness, 1 in SEO Health.
  • Why: an assistant can quote your answer to a customer's question only if one of your pages states it.

Question coverage (note)

Reports how many headings are phrased as questions. Whether the answers under them are good is judged by the graded checks.

Extractable facts (note)

Reports the prices, opening hours and areas served that your pages state as text. Absent categories are not a failing. If your pricing is public, stating it as text or as a "from" price lets assistants answer cost questions.

Thin commercial pages (note)

Lists service, product, pricing and location pages with fewer than 150 words of their own, not counting blocks repeated across the site. Substance is judged by the graded checks.

Summary or key facts near the top (note)

Full scans only. Reports whether long product and service pages open with a short summary or a list of key facts.

Authority

Independent reviews linked

Checks whether your site links to your profile on an independent review platform, such as Google Business Profile, Trustpilot, Yelp, Tripadvisor, G2, Capterra, Feefo or Reviews.io.

  • Passes when it does.
  • Otherwise the check warns and keeps 70% of its points.
  • Ratings marked up on your own pages are neither rewarded nor penalised.
  • Weight: 2 in Site readiness, 0 in SEO Health.
  • Why: reviews on an independent platform are a reputation a reader or an assistant can check, and linking them from the site makes them easy to find. Google does not show review stars a business publishes about itself.

Profiles elsewhere that confirm who you are

Checks the profiles your site links to elsewhere, in your structured data's sameAs or as links on your pages. Review platforms count.

  • Passes with three or more profiles that open, plus a Google Business Profile when customers visit you in person.
  • Warns with one or two profiles, a missing Business Profile for a place people visit, profile links that no longer open, or no profiles at all.
  • Profiles that block automated requests are listed as not checked.
  • A Wikidata entry is noted when found but not required.
  • The free site check judges the links alone.
  • Weight: 4 in Site readiness, 2 in SEO Health.
  • Why: Google documents sameAs as how a business points to its profiles elsewhere, which confirm who it is. A Business Profile is how a place people visit appears in Maps and local results. Source: Organization structured data.

Author and expertise signals

Applies only to article and documentation pages. Checks for a visible, named author.

  • Passes when half or more of those pages have one.
  • Warns when some do.
  • Fails when none does.
  • Author markup is a plus, not required. State qualifications on author bio pages.
  • Weight: 3 in Site readiness, 2 in SEO Health.
  • Why: Google's guidance on helpful content asks whether it is clear who wrote a piece, through bylines and author pages. Source: Creating helpful content.

Freshness signals

Checks the freshest believable date across the site: page modified dates and sitemap dates. Sitemap dates that are all the same, or carry the scan date, are build stamps and are ignored.

  • Warns when the freshest date is over 270 days old, or no date is found.
  • Fails when it is over 540 days old.
  • Do not bump dates without changing the page.
  • Weight: 3 in Site readiness, 4 in SEO Health.
  • Why: Google documents how it reads a page's dates and asks that they be accurate; dates bumped without a real change are not a signal. Source: Publication dates.

Original data and named sources

Applies only to sites with article, guide or documentation pages. Checks for figures of your own, or named sources for the figures you quote. This check never fails.

  • Weight: 2 in Site readiness, 2 in SEO Health.
  • Why: the GEO study (Aggarwal et al., KDD 2024) found that adding statistics and citations of named sources raised how visible a page was in generative-engine answers. Source: GEO: Generative Engine Optimization.

Sites that mention you in tracked AI answers (note)

Reports how many other sites name you in the AI answers Cleotic tracks over the last 90 days.

  • Why it changes no score: it shows where assistants learn about you, but nothing on your own site moves it.

Delivery

Server response time

Measures the time from sending a request for your homepage to the first byte of the answer, as the median of several samples. Measurements are taken from western Europe (the Netherlands), so a site far from there reads a little slower.

  • Passes under 800 ms.
  • Warns under 1.8 seconds, and more strongly under 3 seconds.
  • Fails at 3 seconds or more.
  • Weight: 3 in Site readiness, 6 in SEO Health.
  • Why: Google documents that a server that answers slowly gets crawled less: the crawler slows down to avoid overloading it. Source: Managing crawl budget.

Image descriptions

Checks that meaningful images carry alt text that describes them, not placeholder text.

  • Fails when a quarter or more of meaningful images lack one.
  • Warns below that.
  • Weight: 2 in Site readiness, 3 in SEO Health.
  • Why: Google documents that it uses alt text to understand what an image shows; without it the image says nothing to anything reading the page as text. Source: Google Images best practices.

Image loading

Checks that the first image on the homepage loads straight away and that images declare their width and height.

  • Warns when the first image is lazy-loaded, or more than half the images have no dimensions. This check never fails.
  • Weight: 0 in Site readiness, 2 in SEO Health.
  • Why: a lazy-loaded first image delays the largest paint, and an image without dimensions shifts the layout as it loads. Both are Core Web Vitals measures. Source: Optimize Cumulative Layout Shift.

Payload size and compression

Checks the size of the homepage HTML and whether it is compressed.

  • Fails over 1 MB.
  • Warns over 400 KB, or when the HTML is served uncompressed.
  • Weight: 1 in Site readiness, 4 in SEO Health.
  • Why: Google documents that its crawler reads an HTML file only up to a size limit. Other fetchers have limits too, not always published. Source: Googlebot.

Cache validation headers

Checks for an ETag or Last-Modified header on HTML responses. A Cache-Control policy is not required, and no-store is never penalised.

  • Passes with either header.
  • Warns with neither.
  • Weight: 0 in Site readiness, 2 in SEO Health.
  • Why: Google documents that its crawlers support ETag and Last-Modified, which let them check whether a page changed without downloading it again. Source: Google Search Central: HTTP caching.

Mobile viewport

Checks for a viewport meta tag.

  • Fails when it is missing.
  • Weight: 0 in Site readiness, 3 in SEO Health.
  • Why: Google indexes the mobile version of a page, and a page without a viewport renders at desktop width on a phone. Source: Mobile-first indexing.

Updated