Skip to content
pageinspection

Every check, all 82 in plain language

No black box. This is the full checklist your site gets graded against, plus how the crawl produces it. It replaces the spreadsheet, the crawler and the three browser tabs you had open next to them.

82
checks
10
categories
4
questions
2,000
pages in a single crawl

A list of 82 problems is not a plan

Anyone can hand you a checklist. The work is knowing which of these actually apply to your site, which ones are costing you traffic today versus which are hygiene, and what to do about each one. We publish the full list because there is nothing to hide in it. The product's job is the triage and the fix, not the list.

How the audit runs

One crawl produces everything below. This is the machinery behind the list, not a feature tour.

Crawl & visualize

Crawl up to 2,000 pages in a single run and explore your internal link structure as an interactive tree, a sortable table, or a ranked list of issues.

Page inspector + Lighthouse

Drill into any URL: meta & Open Graph tags, structured data, image alt coverage, and full Core Web Vitals via PageSpeed Insights.

Export anything

Download your full audit as CSV, JSON, or a branded, print-ready PDF report to share with clients and teams.

Four questions, in order

Every check belongs to one of these. They run as a sequence, not a menu: a page a crawler cannot reach never gets judged on whether it earns the click.

Four sequential gates: Crawlable with 24 checks, then Compelling with 34, then Quotable with 16, then Trusted with 8. A page that fails an earlier gate never reaches the later ones.Crawlable24 checksCompelling34 checksQuotable16 checksTrusted8 checks

Crawlable

Can search engines reach every page?

24

The plumbing. If a crawler can't reach a page, nothing else on it matters.

Compelling

Does each page earn the click?

34

Titles, markup and previews: what turns an impression into a visit.

Quotable

Will AI and answer engines cite you?

16

The new layer. Whether assistants and answer boxes can read, trust and quote you.

Trusted

Is it secure and fast enough to rank?

8

Transport security, headers and delivery: the hygiene that quietly caps results.

Critical, Important, Refinement: what each one means

Every check carries one of these, and you can filter by them below. Here is what they actually mean, and how the catalogue splits between them.

Critical: 19 of 82 checks. Important: 45 of 82 checks. Refinement: 18 of 82 checksCritical19 of 82Important45 of 82Refinement18 of 82

Critical

19 of 82

A failure here can stop a page being found, indexed, or read at all. Nothing else you do to that page matters until it is fixed, so these go to the top of the list regardless of how small the change is.

Important

45 of 82

The page still works, but you are competing at a measurable disadvantage against anyone who got it right. This is ordinary remediation work: worth planning, not worth panicking about.

Refinement

18 of 82

Polish. Worth doing once the two above are clear, and not worth blocking a release for. Some of these matter more in aggregate than individually.

The 82 checks in full

Grouped by the four questions. Search it, filter it by area or severity, or read it end to end.

Showing 82 of 82

All 82 checks by category: Technical SEO 13, Architecture 7, International 4, On-page SEO 11, Content quality 7, Structured data 11, Social sharing 5, AI search (GEO) 9, Answer engines (AEO) 7, Security & performance 8. Each bar links to that category.Technical SEO13Architecture7International4On-page SEO11Content quality7Structured data11Social sharing5AI search (GEO)9Answer engines (AEO)7Security & performance8CrawlableCompellingQuotableTrusted

Crawlable

Can search engines reach every page?

The plumbing. If a crawler can't reach a page, nothing else on it matters.

24 checks across 3 categories

01

Technical SEO Health

The fundamentals every crawler checks first

Titles, canonicals, robots rules, redirects and status codes, resolved the way a search engine resolves them, not the way your CMS reports them.

13 checks in this category

  • Page titles

    Critical

    Every page has a title, at a length that survives the results page, saying what the page actually is.

  • Each page ships its own description, so the search engine isn't left to write your snippet for you.

  • One clear H1 per page and no skipped levels, the outline crawlers read to understand your structure.

  • Every page declares its preferred URL, so duplicates don't split ranking authority between them.

  • No stray noindex or nofollow quietly hiding pages you spent months building.

  • Links that hop through two or three redirects before landing. Every hop leaks authority and speed.

  • URL structure

    Refinement

    Readable, lowercase, meaningful paths instead of query-string soup.

  • Crawlability

    Critical

    Whether a bot can genuinely walk from your homepage to this page without hitting a wall.

  • Indexability

    Critical

    Which pages are eligible to appear at all, after robots.txt, meta robots and canonicals are resolved together.

  • Every page has inbound links from elsewhere on the site. Orphan pages don't rank.

  • Click depth

    Important

    How many clicks from the homepage each page sits. Past three or four, crawlers stop coming back.

  • Your important pages are in the XML sitemap, and the sitemap isn't listing pages you've blocked.

  • 404s, 500s and soft errors reached from your own links, with the exact page that links to each.

02

Website Architecture

How your pages connect, and what gets buried

We map the full internal link graph, then show you the hubs, the clusters, the dead ends, and everything sitting too deep to be crawled often.

7 checks in this category

  • Crawl paths

    Critical

    The routes a bot takes through your site, and precisely where they dead-end.

  • URL hygiene

    Important

    Session IDs, tracking parameters and mixed casing, the quiet factory for duplicate URLs.

  • Pagination

    Important

    Paginated series linked so crawlers can walk them and don't treat every page as a duplicate.

  • One canonical form for trailing slashes, instead of two live URLs for every page on the site.

  • The whole graph, rendered as a tree, a force graph, or a sortable table, hubs, clusters and isolated pockets.

  • Inbound and outbound links per URL, so you can see what each page supports and what supports it.

  • Crawl depth

    Important

    The depth distribution across your site, and what is buried too deep to earn traffic.

03

International SEO

Serving the right language to the right visitor

hreflang is easy to get subtly wrong and hard to notice. We verify the declarations are reciprocal, complete, and match what the page actually says.

4 checks in this category

  • Language and region variants declared, and declared reciprocally, which is where most setups break.

  • x-default

    Refinement

    A declared fallback for visitors whose language you don't serve.

  • The language you declare matches the language actually on the page.

  • A lang attribute on the html element, read by screen readers, translators and crawlers alike.

Compelling

Does each page earn the click?

Titles, markup and previews: what turns an impression into a visit.

34 checks across 4 categories

04

On-Page SEO Optimization

Whether each page makes its own case

Unique titles and descriptions, enough substance to rank, honest anchor text, and the trust signals (authors, an About page) that raters and models both look for.

11 checks in this category

  • Pages carrying too little unique text to compete for anything meaningful.

  • Unique titles

    Important

    No two pages sharing a title and cannibalising each other in the results.

  • Templated descriptions repeated across dozens of pages, a visible signal of low effort.

  • Exact and near-identical pages competing with each other instead of being consolidated.

  • Images with no alt attribute, an accessibility failure and a wasted relevance signal.

  • Width and height declared, so your layout doesn't jump while images load.

  • "Click here" and bare URLs tell a search engine nothing about what they point at.

  • Whether your links point where they are genuinely relevant, not just where the template put them.

  • Outbound links

    Refinement

    Pages citing nothing external. Classic and AI search both favour well-sourced content.

  • A substantive About page, one of the strongest trust signals available to both human raters and models.

  • Named, attributable authors on content pages. Anonymous pages struggle on experience and authority.

05

Content Quality

Substance, duplication and internal authority

Which pages are genuinely their own, which are near-duplicates competing with each other, and which ones your own link graph is actually promoting.

7 checks in this category

  • How much of each page is genuinely its own, versus boilerplate repeated site-wide.

  • Exact and near-duplicate clusters, grouped so you can see at a glance what to merge or canonicalise.

  • Thin content

    Important

    Pages too light to satisfy any real search intent, listed worst-first.

  • Whether your heading outline tells a coherent story when read on its own.

  • Sentence and paragraph density, a proxy for how easily people and models parse the page.

  • Which pages your own link graph actually promotes, and whether those are the pages that matter to you.

  • Whether related content is grouped into recognisable topical clusters or scattered across the site.

06

Structured Data Validation

Markup that parses, validates and qualifies

We do not just detect schema. We validate it. Required fields, author and publisher entities, and syntax, across every type from Article to LocalBusiness.

11 checks in this category

  • JSON-LD

    Critical

    Machine-readable markup present at all, the format Google and AI crawlers both prefer.

  • FAQ Schema

    Important

    FAQPage markup on the pages that genuinely answer questions.

  • HowTo Schema

    Important

    Step-by-step content marked up as HowTo, so it can surface as a guided result.

  • Products carrying the price, availability and review fields shopping surfaces require.

  • Editorial content marked up with headline, dates and author.

  • Who publishes this site, the foundation of entity recognition for search engines and language models.

  • Address, opening hours and geo data for anything with a physical presence.

  • Passages explicitly marked as safe for a voice assistant to read aloud.

  • Author and publisher modelled as real entities, not left as bare strings.

  • Each type carrying the properties Google actually requires, not just the convenient ones.

  • Markup that parses. Broken JSON-LD is invisible JSON-LD, and nothing tells you.

07

Social Sharing Optimization

What people see before they see your site

Open Graph and Twitter Card coverage, share images that actually resolve, and canonical agreement so every share consolidates to one URL.

5 checks in this category

  • og:title, og:description and og:image. The first thing every share, DM and AI preview reads.

  • Twitter Cards

    Refinement

    Card markup so your links unfurl as rich previews instead of bare URLs.

  • Social images

    Important

    A share image that exists, resolves, and is sized so it is not cropped into nonsense.

  • Share copy written deliberately, rather than truncated meta text.

  • og:url and your canonical agree, so every share consolidates to the same URL.

Quotable

Will AI and answer engines cite you?

The new layer. Whether assistants and answer boxes can read, trust and quote you.

16 checks across 2 categories

09

Answer Engine Optimization

Turning pages into the answer, not a result

Where you have snippet and FAQ opportunities sitting unclaimed, and whether your answers are structured tightly enough to be lifted verbatim.

7 checks in this category

  • Pages that answer questions but carry no FAQ markup, the cheapest snippet win on your site.

  • Procedural content that could be a step-by-step result, sitting unmarked.

  • A concise, standalone answer an engine can lift verbatim without editing.

  • Headings phrased the way people actually search, rather than the way your org chart talks.

  • Passages short and clean enough for a voice assistant to read out loud.

  • Lists, tables and definitions, the shapes answer boxes reach for before prose.

  • A composite read on whether this page could realistically take position zero.

Trusted

Is it secure and fast enough to rank?

Transport security, headers and delivery: the hygiene that quietly caps results.

8 checks across 1 category

10

Security & Performance

The hygiene that quietly caps your ceiling

TLS, security headers, compression and caching, read straight off the response, plus any mixed content breaking the padlock.

8 checks in this category

  • HTTPS usage

    Critical

    Every page and asset served over TLS. Plain HTTP costs you both ranking and trust.

  • HSTS

    Important

    Strict-Transport-Security, so browsers never even attempt the insecure route.

  • CSP headers

    Important

    A Content-Security-Policy bounding what is allowed to execute on your pages.

  • Protection against your pages being framed and clickjacked.

  • nosniff set, so browsers stop guessing at content types.

  • Compression

    Important

    Gzip or Brotli on text responses, the cheapest speed win available.

  • Cache-Control

    Refinement

    Caching headers that keep repeat visits and repeat crawls cheap.

  • Secure pages pulling in insecure assets, which breaks the padlock and the trust that comes with it.

From a failing check to a shipped fix

The list above is the diagnosis. What you actually work from is a prioritized report: every failing check with the evidence behind it and the list of URLs it affects, so nobody has to re-investigate what the crawler already knew.

Each one exports as a markdown brief or an agent prompt, with the acceptance criteria written in. Then you re-crawl, and the same check either passes or it does not. That last step is the difference between shipping a fix and hoping you did.

See what a fix brief contains
A failing canonical-tag check, its evidence and affected URLs, the exported brief handed to an agent, and a re-crawl that turns the same check green.Canonical tag missingCritical · Technical SEOfailsEvidenceno rel=canonical in <head>Affected URLs/blog/post-a/blog/post-b+12 more14 of 2,000 pagesfix-canonical-tags.mdto your agentCanonical tag presentpassesre-crawl

What this audit doesn't cover

A page selling completeness should say where it stops. These are out of scope, not coming next week.

Backlinks and off-page

No link intersect, no referring-domain counts, no disavow tooling. Everything here is measured on pages we can fetch from your own site.

Keyword volume and rank tracking

We do not tell you what to write about or where you sit for a term this week. The audit is about whether a page can be reached, understood and quoted, not what it targets.

Competitor comparison

No side-by-side scoring against another domain. You can of course run the audit against a competitor URL, but nothing here diffs the two for you.

Analytics and conversion

No traffic, revenue or funnel data. The audit never sees your analytics, and a check passing is not a promise about outcomes.

If you need any of these, you need a different tool alongside this one, and we would rather say so here than after you have signed up.

Questions about running it

Crawl limits, where the performance data comes from, and how often this is worth repeating.

The crawl stops at the ceiling rather than failing, so you get a complete audit of what it reached. Depth and concurrency are configurable, so the usual approach on a large site is to scope a run: start from a section rather than the homepage, or cap the depth, and audit the parts that matter in separate passes instead of trying to swallow everything at once.

Internal HTML pages that we actually fetch. External links are recorded as link targets so the graph is complete, but they are never fetched for their markup and do not consume the budget. Assets like images, scripts and stylesheets are not pages either.

Core Web Vitals come from the PageSpeed Insights API, requested for the URL you are inspecting at the moment you ask for it. That means it is a lab measurement taken then and there, not a rolling average of your real visitors, so treat it as a diagnostic signal rather than a report on field performance.

After anything that changes templates, routing or rendering, which is the moment regressions get introduced, and on a slow cadence otherwise. The more useful habit is targeted: after you ship a fix, re-run and confirm that specific check flipped to passing. A quarterly full crawl catches the drift nobody meant to cause.

A staging site is fine as long as it is publicly reachable and not disallowed in its own robots.txt. Anything behind a login or HTTP auth is not, because the crawler fetches pages as an anonymous visitor and holds no credentials for your site. If a page needs a session to render its content, we see what a signed-out visitor sees, which is also what a search engine sees.

They have to. The AI search and answer engine categories barely existed a couple of years ago, and llms.txt is an emerging convention rather than a settled standard. The catalogue on this page is the current one, and it grows as the surfaces do, which is part of why it is published in full rather than summarised.

See which of these your site fails

The free preview shows headline numbers for a single page. A free account crawls the whole site and runs every check across it, PDF export included.

Try it free

No signup required. Each free search audits one page, paste any URL to see it in action.