Technical SEO Audit

Technical SEO Audit Services, Free in Five Business Days

seothumb.com runs a full website crawl audit of your site, reconciles it against Search Console and your server logs, and returns a twelve-page document with every finding scored by engineering effort against the revenue it protects. It costs nothing and does not require a call.

  • 12 pages

    Audit length

  • 5 business days

    Turnaround

  • $0, no call required

    Cost

Two colleagues reviewing code together on a monitor

What the 12-page technical SEO audit actually contains

Every audit ships as the same twelve numbered pages, so you can hold it against anything else you have been sent and compare like for like. Nothing is padded with a tool's default export.

  1. Findings summary — the three issues capping the site, on one page, in language a non-technical stakeholder can act on.
  2. Crawl overview — URLs discovered, status code distribution, depth histogram, templates identified and sized.
  3. Indexation reconciliation — the crawlable set joined against the Search Console Pages report and bucketed four ways.
  4. Index bloat — what is indexed that should not be, and the mechanism producing it.
  5. Canonical control — self-reference rate, conflicts, loops and cross-template contradictions.
  6. Redirects — chains, loops, soft 404s and internal links still pointing at redirected URLs.
  7. Core Web Vitals field review — CrUX data split by template and device, with the failing LCP element named.
  8. Rendering — what exists in the raw HTML versus after JavaScript execution, and which content Google may never see.
  9. Internal link graph — depth, inlink distribution, orphan list, and pages absorbing links they do not earn.
  10. Log-file summary — Googlebot hit frequency by URL class, when access logs are available.
  11. Structured data — validation errors, missing types and markup that contradicts the visible page.
  12. Prioritized fix list — every finding scored, sequenced and assigned to engineering, CMS or content.

The raw crawl export ships alongside it, so nothing in the document is a claim you cannot check yourself.

How the website crawl audit is run

The crawl is run in Screaming Frog, in database storage mode so a large catalog does not get truncated at an arbitrary URL ceiling, with no depth limit. It runs twice.

The first pass uses the Googlebot smartphone user agent with JavaScript rendering enabled, an AJAX timeout long enough for hydration to finish, and a mobile viewport. The second pass is text-only. Diffing the two shows exactly which links, canonicals, headings and body copy exist only after rendering — the content a crawler may discount or miss entirely on a slow render queue.

The first pass respects robots.txt, because that is the site Google sees. A third pass ignores robots.txt deliberately, to inventory what is being hidden. Directories blocked in robots.txt are a common place for indexed-but-uncrawlable URLs to accumulate, and you cannot count them without looking inside.

The Search Console, Ahrefs and PageSpeed Insights APIs are connected to the crawl, so every row carries clicks, impressions, referring domains and field Core Web Vitals next to its status code. Custom extraction pulls canonical tags, hreflang, pagination markup, schema blocks and any CMS identifier that lets URLs be grouped by template.

Request rate is capped and scheduled off-peak, and we ask for an allowlist rather than crawling from an unknown IP. XML sitemaps are crawled separately in list mode, then compared against the discovered set. Google's SEO starter guide describes the crawl-to-index path this configuration is designed to reproduce.

Log-file analysis and what Googlebot hit frequency reveals

A crawl tells you what Google could reach. Access logs tell you what it actually did. We ask for thirty days of raw logs — nginx or Apache, or a Logpush or CloudFront export if a CDN terminates the requests — and verify Googlebot by reverse DNS lookup rather than trusting the user-agent string, because a meaningful share of self-declared Googlebot traffic is not Google.

Verified hits are then grouped by URL class rather than by URL, and the distribution is the finding. When parameter and faceted URLs absorb the majority of requests, the pages that earn money are being crawled on a slower cycle than the pages that do not exist for buyers. Other patterns worth naming:

  • Days since last hit on revenue-carrying templates, compared against templates nobody needs crawled.
  • Discovery lag between publish time and first crawl, which is the real ceiling on how fast new content can rank.
  • Status codes served to Googlebot specifically — intermittent 5xx or 429 responses under crawl load often never appear in human-facing monitoring.
  • URLs Googlebot requests that the crawl never found, which is one of the two reliable ways to detect orphan pages.
  • Cache and 304 behavior, which shapes how much of the crawl allowance is spent re-fetching unchanged pages.

Without logs the audit still ships; the Crawl Stats report in Search Console stands in, at far coarser resolution. Logs are the single highest-value input you can hand over.

Index bloat, coverage reconciliation, and canonical conflicts

The Search Console Pages report is exported in full, per status and per reason, then joined against the crawl on URL — in a spreadsheet for a small site, in BigQuery once the row count makes that impractical. Every URL lands in one of four buckets: indexed and should be, indexed and should not be, crawlable but not indexed, and neither. Only the join produces those buckets. The site: operator does not, and neither report alone does.

Where Google's chosen canonical is in doubt, URLs are sampled through the URL Inspection API, which returns the canonical Google selected next to the one you declared. A gap between those two is the finding — not the tag you wrote.

Canonical and redirect conflicts that a website crawl audit surfaces repeatedly:

  • Paginated pages canonicalizing to page one, which quietly removes deep products from the index.
  • Canonical tags pointing at a URL that redirects, or at a noindexed URL.
  • Faceted URLs self-canonicalizing while also sitting in the sitemap and being internally linked.
  • 302s standing in for permanent moves, and redirect chains from http to https to a trailing-slash variant.
  • Out-of-stock and discontinued URLs returning soft 404s instead of consolidating their links.

The remediation order matters more than the tags. Blocking a bloated directory in robots.txt before those URLs are deindexed means Google never re-reads the noindex, and the URLs stay in the index indefinitely. The fix list sequences that correctly. Template-level implementation is covered under on-page SEO.

Core Web Vitals field data, and why lab scores mislead

The Core Web Vitals review is built on CrUX field data — the 75th percentile of real Chrome users over a rolling 28-day window — not on a synthetic run. Where a URL has too little traffic for URL-level data, we report origin-level and template-level groupings and say which is which, rather than presenting one as the other.

A lab score is a single load, on a throttled synthetic connection, from a cold cache, on a machine that is not your customer's phone. It usually runs without a consent banner, without a logged-in session, and often without the third-party tags that fire for real visitors. That is why a green lab result and a failing field assessment sit side by side so often. The pass mark is 2.5 seconds for Largest Contentful Paint, 200 milliseconds for Interaction to Next Paint and 0.1 for Cumulative Layout Shift, measured at the 75th percentile, and the definitions are documented in Google's Core Web Vitals reference.

The review names the cause, not the score. LCP is broken into its four parts — time to first byte, resource load delay, resource load duration and element render delay — so the fix is specific: a preload, a server change, an image format, or a hero that only renders after hydration. INP is diagnosed from real interactions, which a lab test cannot generate at all. Phone and desktop are always reported separately.

Internal link graph and orphan pages

Internal linking is where sites give away rankings without noticing, because nothing visibly breaks. The audit builds the link graph from the crawl and reads three things from it.

Depth. How many clicks from the homepage each template sits at, plotted as a distribution. Commercially important URLs buried at depth six behind pagination are being told, structurally, that they do not matter.

Inlink distribution against value. Which URLs receive the most internal links, ranked beside what those URLs are worth. A privacy policy sitting in the footer of 40,000 pages routinely outranks a money category on internal links received. Navigation and footer links are counted separately from in-body contextual links, because they do not carry the same weight or the same anchor variety.

Orphans. A URL is flagged as orphaned when it appears in the XML sitemap, GA4, Search Console or the server logs but has zero internal inlinks in the crawl. That cross-source join is the only reliable detection method. Orphans accumulate from retired categories, campaign landing pages, CMS entries unpublished from navigation, and product URLs that fell out of a feed.

The output is a set of template-level link rules, not a list of individual pages to link by hand — one deploy that changes a related-items module or a breadcrumb reaches far more URLs than any manual pass. Catalog-specific handling is covered under ecommerce SEO.

How the fix list is scored, and what happens after delivery

Every finding carries three values. Effort, banded in engineering days: under a day, one to three, three to ten, ten or more. Confidence, stated plainly, because a canonical conflict is a certainty and a rendering change is an expectation. Forecast impact, expressed as the revenue attached to the URL cluster the fix affects, calculated from your GA4 conversion data, your average order or contract value and your close rate — your numbers, not a benchmark.

Those three place each item in one of four lanes: ship this sprint, schedule, batch into the next planned release, or park with a note explaining why it is not worth doing. Each line names an owner — engineering, CMS or content — so the list can go straight into a backlog without translation.

The audit arrives by email as a PDF with the crawl export attached. It is free, no call is required to receive it, and it is not held back behind a proposal. If your developers implement it themselves, that is a good outcome, and we will answer questions by email while they do.

If the crawl comes back clean and there is nothing on your site worth a retainer, the summary page will say exactly that. A free audit that manufactures urgency is worth less than no audit at all. Where the work does justify ongoing help, scope and cost are published on the pricing page and the wider engagement model is set out under SEO services.

FAQ

Questions about technical seo audit

  • What access do you need to run the audit?

    Read-only Search Console and GA4 access, and permission to crawl from a known IP at a capped request rate. Thirty days of raw access logs, or a CDN log export, make the audit considerably sharper but are not required. No CMS login, no hosting credentials and no code access are needed to produce the document.

  • Why not just run Screaming Frog ourselves?

    A tool produces findings; it cannot rank them. The value here is the joins — crawl against Search Console coverage, crawl against verified Googlebot log hits, rendered HTML against raw HTML, field Core Web Vitals against template — and then scoring each result against the revenue the affected URLs carry. That judgment is the deliverable. The crawl is raw input.

  • Can our own developers implement the fix list?

    Yes, and the document is written for that. Each item names an owner, an effort band and enough technical specificity for an engineer to open a ticket without a follow-up call. We answer implementation questions by email either way. Plenty of technical debt is best cleared by the team that already knows the codebase.

  • How often should a technical SEO audit be repeated?

    A full re-audit once a year suits most sites, and sooner after a replatform, a migration, a navigation rebuild or a CMS upgrade — those four events create more indexation damage than anything else. Between full audits, reviewing the Search Console Pages report and field Core Web Vitals monthly catches drift early enough to fix it cheaply.

Next step

Send your domain, get the audit in five business days

No call, no proposal deck, no cost. Send your URL through the [contact form](/contact) and you get the twelve-page technical SEO audit, the crawl export and the scored fix list inside five business days. If the site is already in good shape, the summary will say so and you can stop there.

Get my free technical SEO audit

Five business day turnaround. No call required to receive it.