Every check, all 82 in plain language
No black box. This is the full checklist your site gets graded against, plus how the crawl produces it. It replaces the spreadsheet, the crawler and the three browser tabs you had open next to them.
- 82
- checks
- 10
- categories
- 4
- questions
- 2,000
- pages in a single crawl
A list of 82 problems is not a plan
Anyone can hand you a checklist. The work is knowing which of these actually apply to your site, which ones are costing you traffic today versus which are hygiene, and what to do about each one. We publish the full list because there is nothing to hide in it. The product's job is the triage and the fix, not the list.
How the audit runs
One crawl produces everything below. This is the machinery behind the list, not a feature tour.
Crawl & visualize
Crawl up to 2,000 pages in a single run and explore your internal link structure as an interactive tree, a sortable table, or a ranked list of issues.
Page inspector + Lighthouse
Drill into any URL: meta & Open Graph tags, structured data, image alt coverage, and full Core Web Vitals via PageSpeed Insights.
Export anything
Download your full audit as CSV, JSON, or a branded, print-ready PDF report to share with clients and teams.
Four questions, in order
Every check belongs to one of these. They run as a sequence, not a menu: a page a crawler cannot reach never gets judged on whether it earns the click.
Crawlable
Can search engines reach every page?
The plumbing. If a crawler can't reach a page, nothing else on it matters.
Compelling
Does each page earn the click?
Titles, markup and previews: what turns an impression into a visit.
Quotable
Will AI and answer engines cite you?
The new layer. Whether assistants and answer boxes can read, trust and quote you.
Trusted
Is it secure and fast enough to rank?
Transport security, headers and delivery: the hygiene that quietly caps results.
Critical, Important, Refinement: what each one means
Every check carries one of these, and you can filter by them below. Here is what they actually mean, and how the catalogue splits between them.
Critical
19 of 82A failure here can stop a page being found, indexed, or read at all. Nothing else you do to that page matters until it is fixed, so these go to the top of the list regardless of how small the change is.
Important
45 of 82The page still works, but you are competing at a measurable disadvantage against anyone who got it right. This is ordinary remediation work: worth planning, not worth panicking about.
Refinement
18 of 82Polish. Worth doing once the two above are clear, and not worth blocking a release for. Some of these matter more in aggregate than individually.
The 82 checks in full
Grouped by the four questions. Search it, filter it by area or severity, or read it end to end.
Showing 82 of 82
Crawlable
Can search engines reach every page?
The plumbing. If a crawler can't reach a page, nothing else on it matters.
24 checks across 3 categories
Technical SEO Health
The fundamentals every crawler checks first
Titles, canonicals, robots rules, redirects and status codes, resolved the way a search engine resolves them, not the way your CMS reports them.
13 checks in this category
Page titles
CriticalEvery page has a title, at a length that survives the results page, saying what the page actually is.
Meta descriptions
ImportantEach page ships its own description, so the search engine isn't left to write your snippet for you.
Heading hierarchy (H1 to H6)
ImportantOne clear H1 per page and no skipped levels, the outline crawlers read to understand your structure.
Canonical tags
CriticalEvery page declares its preferred URL, so duplicates don't split ranking authority between them.
Robots directives
CriticalNo stray noindex or nofollow quietly hiding pages you spent months building.
Redirect chains
ImportantLinks that hop through two or three redirects before landing. Every hop leaks authority and speed.
URL structure
RefinementReadable, lowercase, meaningful paths instead of query-string soup.
Crawlability
CriticalWhether a bot can genuinely walk from your homepage to this page without hitting a wall.
Indexability
CriticalWhich pages are eligible to appear at all, after robots.txt, meta robots and canonicals are resolved together.
Internal linking
CriticalEvery page has inbound links from elsewhere on the site. Orphan pages don't rank.
Click depth
ImportantHow many clicks from the homepage each page sits. Past three or four, crawlers stop coming back.
Sitemap inclusion
ImportantYour important pages are in the XML sitemap, and the sitemap isn't listing pages you've blocked.
HTTP status codes
Critical404s, 500s and soft errors reached from your own links, with the exact page that links to each.
Website Architecture
How your pages connect, and what gets buried
We map the full internal link graph, then show you the hubs, the clusters, the dead ends, and everything sitting too deep to be crawled often.
7 checks in this category
Crawl paths
CriticalThe routes a bot takes through your site, and precisely where they dead-end.
URL hygiene
ImportantSession IDs, tracking parameters and mixed casing, the quiet factory for duplicate URLs.
Pagination
ImportantPaginated series linked so crawlers can walk them and don't treat every page as a duplicate.
Slash consistency
ImportantOne canonical form for trailing slashes, instead of two live URLs for every page on the site.
Internal link graph
RefinementThe whole graph, rendered as a tree, a force graph, or a sortable table, hubs, clusters and isolated pockets.
Page relationships
RefinementInbound and outbound links per URL, so you can see what each page supports and what supports it.
Crawl depth
ImportantThe depth distribution across your site, and what is buried too deep to earn traffic.
International SEO
Serving the right language to the right visitor
hreflang is easy to get subtly wrong and hard to notice. We verify the declarations are reciprocal, complete, and match what the page actually says.
4 checks in this category
hreflang implementation
ImportantLanguage and region variants declared, and declared reciprocally, which is where most setups break.
x-default
RefinementA declared fallback for visitors whose language you don't serve.
Language consistency
ImportantThe language you declare matches the language actually on the page.
HTML lang attribute
ImportantA lang attribute on the html element, read by screen readers, translators and crawlers alike.
Compelling
Does each page earn the click?
Titles, markup and previews: what turns an impression into a visit.
34 checks across 4 categories
On-Page SEO Optimization
Whether each page makes its own case
Unique titles and descriptions, enough substance to rank, honest anchor text, and the trust signals (authors, an About page) that raters and models both look for.
11 checks in this category
Content length
ImportantPages carrying too little unique text to compete for anything meaningful.
Unique titles
ImportantNo two pages sharing a title and cannibalising each other in the results.
Unique descriptions
RefinementTemplated descriptions repeated across dozens of pages, a visible signal of low effort.
Duplicate content
CriticalExact and near-identical pages competing with each other instead of being consolidated.
Image alt text
ImportantImages with no alt attribute, an accessibility failure and a wasted relevance signal.
Image dimensions
RefinementWidth and height declared, so your layout doesn't jump while images load.
Descriptive anchor text
Important"Click here" and bare URLs tell a search engine nothing about what they point at.
Internal link quality
RefinementWhether your links point where they are genuinely relevant, not just where the template put them.
Outbound links
RefinementPages citing nothing external. Classic and AI search both favour well-sourced content.
About page quality
ImportantA substantive About page, one of the strongest trust signals available to both human raters and models.
Author information
ImportantNamed, attributable authors on content pages. Anonymous pages struggle on experience and authority.
Content Quality
Substance, duplication and internal authority
Which pages are genuinely their own, which are near-duplicates competing with each other, and which ones your own link graph is actually promoting.
7 checks in this category
Content uniqueness
ImportantHow much of each page is genuinely its own, versus boilerplate repeated site-wide.
Duplicate pages
CriticalExact and near-duplicate clusters, grouped so you can see at a glance what to merge or canonicalise.
Thin content
ImportantPages too light to satisfy any real search intent, listed worst-first.
Heading organization
ImportantWhether your heading outline tells a coherent story when read on its own.
Readability signals
RefinementSentence and paragraph density, a proxy for how easily people and models parse the page.
Internal authority
ImportantWhich pages your own link graph actually promotes, and whether those are the pages that matter to you.
Site architecture
RefinementWhether related content is grouped into recognisable topical clusters or scattered across the site.
Structured Data Validation
Markup that parses, validates and qualifies
We do not just detect schema. We validate it. Required fields, author and publisher entities, and syntax, across every type from Article to LocalBusiness.
11 checks in this category
JSON-LD
CriticalMachine-readable markup present at all, the format Google and AI crawlers both prefer.
FAQ Schema
ImportantFAQPage markup on the pages that genuinely answer questions.
HowTo Schema
ImportantStep-by-step content marked up as HowTo, so it can surface as a guided result.
Product Schema
ImportantProducts carrying the price, availability and review fields shopping surfaces require.
Article Schema
ImportantEditorial content marked up with headline, dates and author.
Organization Schema
CriticalWho publishes this site, the foundation of entity recognition for search engines and language models.
LocalBusiness Schema
ImportantAddress, opening hours and geo data for anything with a physical presence.
Speakable Schema
RefinementPassages explicitly marked as safe for a voice assistant to read aloud.
Author & Publisher properties
ImportantAuthor and publisher modelled as real entities, not left as bare strings.
Required schema fields
CriticalEach type carrying the properties Google actually requires, not just the convenient ones.
Schema syntax validation
CriticalMarkup that parses. Broken JSON-LD is invisible JSON-LD, and nothing tells you.
Quotable
Will AI and answer engines cite you?
The new layer. Whether assistants and answer boxes can read, trust and quote you.
16 checks across 2 categories
AI Search, Generative Engine Optimization
Whether AI search can read you at all
llms.txt, AI crawler access, and whether your content survives without JavaScript. Most sites fail here without knowing the category exists.
9 checks in this category
llms.txt
ImportantThe emerging standard that tells AI crawlers what your site is and where the substance lives.
llms-full.txt
RefinementThe expanded manifest: your key content in one clean, model-readable file.
AI crawler accessibility
CriticalWhether GPTBot, ClaudeBot, PerplexityBot and their peers are allowed in. Many sites block them by accident.
AI bot directives
ImportantYour robots.txt and meta rules for AI agents, read back to you explicitly instead of assumed.
Text extractability
CriticalWhether your content survives without JavaScript. If a model can't read it, you don't exist to it.
Content segmentation
ImportantClear, self-contained sections a model can quote without dragging in half the page.
Answer-first content
ImportantThe answer near the top, ahead of the preamble, how models decide what is worth lifting.
Question markup
ImportantQuestions marked up as questions, so the answer beneath them is attributable to you.
AI-friendly structure
ImportantSemantic HTML (real headings, lists and tables) instead of a thousand nested divs.
Answer Engine Optimization
Turning pages into the answer, not a result
Where you have snippet and FAQ opportunities sitting unclaimed, and whether your answers are structured tightly enough to be lifted verbatim.
7 checks in this category
FAQ opportunities
ImportantPages that answer questions but carry no FAQ markup, the cheapest snippet win on your site.
HowTo opportunities
ImportantProcedural content that could be a step-by-step result, sitting unmarked.
Direct answer sections
ImportantA concise, standalone answer an engine can lift verbatim without editing.
Question-based headings
RefinementHeadings phrased the way people actually search, rather than the way your org chart talks.
Speakable content
RefinementPassages short and clean enough for a voice assistant to read out loud.
Structured answers
ImportantLists, tables and definitions, the shapes answer boxes reach for before prose.
Featured snippet readiness
CriticalA composite read on whether this page could realistically take position zero.
Trusted
Is it secure and fast enough to rank?
Transport security, headers and delivery: the hygiene that quietly caps results.
8 checks across 1 category
Security & Performance
The hygiene that quietly caps your ceiling
TLS, security headers, compression and caching, read straight off the response, plus any mixed content breaking the padlock.
8 checks in this category
HTTPS usage
CriticalEvery page and asset served over TLS. Plain HTTP costs you both ranking and trust.
HSTS
ImportantStrict-Transport-Security, so browsers never even attempt the insecure route.
CSP headers
ImportantA Content-Security-Policy bounding what is allowed to execute on your pages.
X-Frame-Options
ImportantProtection against your pages being framed and clickjacked.
X-Content-Type-Options
Refinementnosniff set, so browsers stop guessing at content types.
Compression
ImportantGzip or Brotli on text responses, the cheapest speed win available.
Cache-Control
RefinementCaching headers that keep repeat visits and repeat crawls cheap.
Mixed content
CriticalSecure pages pulling in insecure assets, which breaks the padlock and the trust that comes with it.
From a failing check to a shipped fix
The list above is the diagnosis. What you actually work from is a prioritized report: every failing check with the evidence behind it and the list of URLs it affects, so nobody has to re-investigate what the crawler already knew.
Each one exports as a markdown brief or an agent prompt, with the acceptance criteria written in. Then you re-crawl, and the same check either passes or it does not. That last step is the difference between shipping a fix and hoping you did.
See what a fix brief containsWhat this audit doesn't cover
A page selling completeness should say where it stops. These are out of scope, not coming next week.
Backlinks and off-page
No link intersect, no referring-domain counts, no disavow tooling. Everything here is measured on pages we can fetch from your own site.
Keyword volume and rank tracking
We do not tell you what to write about or where you sit for a term this week. The audit is about whether a page can be reached, understood and quoted, not what it targets.
Competitor comparison
No side-by-side scoring against another domain. You can of course run the audit against a competitor URL, but nothing here diffs the two for you.
Analytics and conversion
No traffic, revenue or funnel data. The audit never sees your analytics, and a check passing is not a promise about outcomes.
If you need any of these, you need a different tool alongside this one, and we would rather say so here than after you have signed up.
Questions about running it
Crawl limits, where the performance data comes from, and how often this is worth repeating.
The crawl stops at the ceiling rather than failing, so you get a complete audit of what it reached. Depth and concurrency are configurable, so the usual approach on a large site is to scope a run: start from a section rather than the homepage, or cap the depth, and audit the parts that matter in separate passes instead of trying to swallow everything at once.
Internal HTML pages that we actually fetch. External links are recorded as link targets so the graph is complete, but they are never fetched for their markup and do not consume the budget. Assets like images, scripts and stylesheets are not pages either.
Core Web Vitals come from the PageSpeed Insights API, requested for the URL you are inspecting at the moment you ask for it. That means it is a lab measurement taken then and there, not a rolling average of your real visitors, so treat it as a diagnostic signal rather than a report on field performance.
After anything that changes templates, routing or rendering, which is the moment regressions get introduced, and on a slow cadence otherwise. The more useful habit is targeted: after you ship a fix, re-run and confirm that specific check flipped to passing. A quarterly full crawl catches the drift nobody meant to cause.
A staging site is fine as long as it is publicly reachable and not disallowed in its own robots.txt. Anything behind a login or HTTP auth is not, because the crawler fetches pages as an anonymous visitor and holds no credentials for your site. If a page needs a session to render its content, we see what a signed-out visitor sees, which is also what a search engine sees.
They have to. The AI search and answer engine categories barely existed a couple of years ago, and llms.txt is an emerging convention rather than a settled standard. The catalogue on this page is the current one, and it grows as the surfaces do, which is part of why it is published in full rather than summarised.
See which of these your site fails
The free preview shows headline numbers for a single page. A free account crawls the whole site and runs every check across it, PDF export included.
No signup required. Each free search audits one page, paste any URL to see it in action.
Social Sharing Optimization
What people see before they see your site
Open Graph and Twitter Card coverage, share images that actually resolve, and canonical agreement so every share consolidates to one URL.
5 checks in this category
Open Graph tags
Importantog:title, og:description and og:image. The first thing every share, DM and AI preview reads.
Twitter Cards
RefinementCard markup so your links unfurl as rich previews instead of bare URLs.
Social images
ImportantA share image that exists, resolves, and is sized so it is not cropped into nonsense.
Social descriptions
RefinementShare copy written deliberately, rather than truncated meta text.
Canonical consistency
Importantog:url and your canonical agree, so every share consolidates to the same URL.