A technical SEO audit is a review of whether a search engine can discover, crawl, render, index and retrieve the URLs that matter on a site, and of how efficiently it does so. The crawl-based checks (robots rules, status codes, canonicals, page speed) are the easy part; any site-audit tool runs them in a click. The audit is the reading that follows: which flagged issues stop Google on the path from discovery to retrieval, which are noise, and what the evidence is for each.
The finding that shows why the two are different things comes up on JavaScript sites more than anywhere else. A category template drops out of the index and the crawl shows nothing wrong with it: 200 status, self-canonical, in the sitemap, linked from the menu. The cause is a bundle that injects the product grid after load, so Googlebot has been indexing the frame around the grid. No checker flags it, because every checker reads the raw file and the raw file is fine. What follows is the path in order, with the Search Console statuses, crawler filters and document limits named, and a report format at the end that a developer can act on.
What a technical SEO audit checks: the six failure states
A technical audit is one question asked six ways: can the engine discover, crawl, render, understand, index and retrieve the pages you need it to? Every issue a crawler flags is a symptom of a failure at one of six points on that path.
- Undiscovered. The engine never finds the URL.
- Uncrawlable. It finds the URL and can’t fetch it.
- Unrenderable. It fetches the file and can’t build the page a user sees.
- Unindexable. It renders the page and doesn’t grant it a place in the index.
- Misunderstood. It indexes the page and reads it as being about the wrong thing, or as a copy of another URL.
- Inefficient. It does all of the above, and spends so much effort on junk URLs that the important ones wait weeks.
Filed against those six states, a long issue list reorganizes itself. Most of the reds turn out to be cosmetic, and a few yellows turn out to be the reason a template is invisible.
Scope: this is the machine-access layer. Whether the pages deserve to rank once indexed — demand, content depth, authority — is a different audit, and this one assumes that question is handled elsewhere.
Before the audit: access, and when to run one
Four things have to be in hand before the reading starts: owner access to Search Console, because the API limits below are per verified property; a month of raw server logs; a full crawl with JavaScript rendering on; and the site’s own list of URLs that carry revenue or leads, because impact is weighted by that list and nothing in a crawl can supply it.
A full audit is due after any migration, redesign or platform change, when the Pages report shows a template moving into an excluded status, and when new content takes weeks to appear in the index. Between those events the work is a recheck, which the FAQ at the end sets out.
Crawlability: robots.txt, crawl budget, crawl traps and orphan pages
Discovery is where the path starts, and it fails more silently than anything downstream, because an undiscovered page throws no error.
Can Googlebot reach the URL
Crawlability is two questions the tools tend to merge. First, can Googlebot reach the URL at all: is it linked from somewhere the crawler travels, present in a sitemap, and outside every Disallow in robots.txt. Second, once reachable, is the crawler’s effort being spent on it or burned elsewhere. A site can be perfectly crawlable in the first sense and starved in the second.
The first-sense failures cluster in a few shapes. A robots.txt rule blocking a directory someone forgot was important, a staging rule that shipped to production, a Disallow written for one path that catches three. And orphan pages: indexable, revenue-carrying URLs with no internal link pointing at them, so the crawler has no path to them.
Crawl budget
CRAWL BUDGET, AS GOOGLE DEFINES IT
Google allots each site a rough amount of crawler time and requests and spends it. Its crawl-budget guidance says the budget matters above roughly a million pages that change weekly, or ten thousand that change daily, or on any site where Search Console shows a large share of URLs as Discovered – currently not indexed. The rewrite of that guidance published on 22 July 2026 states two things in so many words: every site starts on the same conservative default limit, which grows only as the server proves it can take more, and “the crawl capacity limit is shared across all crawlers”, so a heavy AdsBot run is spending Googlebot’s allowance.
Crawl traps
Crawl traps are infinite URL spaces the crawler falls into and never climbs out of: a calendar that generates a “next month” link forever, a faceted navigation that multiplies filters into millions of combinations, session IDs stamped into every URL so one page looks like a thousand. The faceted and parameter explosion does the most damage on larger sites, because beyond hiding pages it drowns the crawler in near-duplicates while the pages that matter wait. Which facets deserve a page is a revenue call that belongs to the ecommerce SEO audit; the concern here is mechanical, whether the explosion is eating the crawl.
Reading Crawl stats three ways
I read crawlability from a full-site crawl segmented by URL type (templates, parameters, facets, pagination), set against Search Console’s Crawl stats report and the sitemap. Crawl stats splits Googlebot’s requests three ways. By purpose: a mature site whose crawl is still mostly “discovery” has a crawler that keeps finding URLs it didn’t know about, and on a site that hasn’t grown, those are traps. By response: the share of requests answered with 3xx, 404 and 5xx is crawl spent on nothing, and after a couple of migrations it often runs into double digits. By Googlebot type: since the July rewrite this is where the shared budget shows, with AdsBot and the image fetcher drawing from the same pool as Googlebot Smartphone. Orphans surface where a URL sits in the sitemap or the logs and has no internal link; traps surface as URL patterns that balloon far past the number of pages the site has. That locates the break. Rebuilding the internal links or rewriting the robots rules is the build work that follows.
Indexability: Search Console index statuses, noindex, canonical, soft 404
Crawling is whether Googlebot can fetch the URL. Indexing is whether Google grants it an independent place in the results. Between the two sits every page that gets fetched, considered and dropped, and no crawl-only check reads that layer.
The statuses and where each fix lives
Search Console names the gap in specific statuses, and each one is a different diagnosis pointing at a different section of this article.
| Search Console status | What it means | Failure state | Where the fix lives |
|---|---|---|---|
Discovered – currently not indexed | Google knows the URL and hasn’t fetched it; the crawler is rationing | Inefficient | Crawlability, performance, logs |
Crawled – currently not indexed | Google fetched it, looked, and declined: quality or duplication | Unindexable | Content audit, canonicalization |
Duplicate, Google chose different canonical than user | Your canonical tag was overruled; Google picked another URL | Misunderstood | URL architecture |
Duplicate without user-selected canonical | No canonical at all on a page Google sees as a copy | Misunderstood | URL architecture |
Indexed, though blocked by robots.txt | Blocked after it was indexed; the noindex can’t be read | Uncrawlable | The trap below |
Excluded by 'noindex' tag | A noindex is doing its job; check it was meant to | Unindexable | Indexability |
Soft 404 | Returns 200; Google read it as empty | Unindexable | HTTP status, rendering |
Page with redirect | The URL redirects; the destination is what gets judged | — | HTTP status, sitemaps |
The mechanisms behind the statuses
A noindex tag doing its job somewhere nobody meant it to, applied site-wide during a redesign and never lifted. A canonical tag pointing the wrong way, telling Google the real page is a duplicate of a weaker one, so the weaker one ranks and the real one is suppressed; Google’s documentation calls the canonical a hint, and when it overrules you the page lands in the third row above. Soft 404s: pages that return 200 OK while showing an empty search result, a “no products found,” a placeholder. Google’s soft-404 handling goes by what the page shows rather than by the status code, so the URL is judged worthless and dropped.
Index bloat and sitemap parity
Two failures only show up when you compare lists. Index bloat: thousands of thin, auto-generated URLs indexed that never should have been. And sitemap-to-index mismatch: the sitemap declares one count, the index holds a smaller one, and the gap is either pages Google refused or pages you never meant to submit. Either direction is a finding.
None of this is visible from what the site intends; you have to read what Google holds. A site: search won’t give a straight count, and on a large site it can be off by an order of magnitude. The URL Inspection API is exact and capped at 2,000 requests a day per property, so on a large site I check class by class, sampled: a bulk index check across Google and Bing, template by template, reading live index status, so the picture is “of this class of page, here is the share that’s indexed,” set against what the site thinks it published. A class where the canonical or noindex on the page disagrees with what’s in the index is where the finding sits.
noindex, a URL that only 301s, a canonical pointing elsewhere.On very large sites the unit of analysis stops being the page and becomes the template, which is the ground of the enterprise SEO audit.
How Googlebot renders JavaScript, and what it never does
Rendering is the step between fetching a file and reading a page, and on JavaScript sites the two can be different documents. When a site is built in React, Vue, Angular or any similar framework, the server often ships a near-empty HTML shell, a <div id="root"> and a bundle of scripts, and the browser runs the scripts, fetches the data and builds the page. That step is hydration; the raw file is an empty frame and the content is poured in afterward.
Google renders JavaScript on a separate pass. It fetches the raw HTML first, queues the page for rendering, and comes back later to run the scripts. In that gap, or if a script errors, or if the content depends on an event the crawler never fires, the rendered page can come back missing the things you need indexed. Google’s JavaScript documentation says the renderer doesn’t scroll or click; practitioners’ tests show it renders with a tall viewport, which is why lazy-loading built on IntersectionObserver usually fires while anything behind a scroll event or a “Load more” button never exists for it.
The failures I look for: content in the rendered DOM and absent from the raw HTML. Internal links injected by JavaScript with no <a href> underneath, so the crawler sees no path onward and discovery dies one level deep. Metadata set by script after load (titles, canonicals, meta robots), where the crawler may read one thing and the user another; Google’s guidance says to avoid changing the canonical with JavaScript at all. Lazy-loaded content and infinite scroll with no crawlable pagination beneath them, so everything past the first screen is unreachable.
I catch it by crawling the page twice, once reading raw HTML and once with rendering on, and diffing the two. In Screaming Frog that is JavaScript rendering mode with both the original and rendered HTML stored, and the JavaScript tab does the diff: Contains JavaScript Content, Contains JavaScript Links, Canonical Only in Rendered HTML, Noindex Only in Original HTML, Page Title Updated by JavaScript. Each filter is one of the failures above, counted per template. For a single URL the same check is Search Console’s URL Inspection: view crawled page, HTML tab against the screenshot. If the content grid is in the screenshot and absent from the HTML, that is the empty frame. Where the gap is links, the follow-up question is whether the deeper pages are reachable any other way.
Where duplicate URLs come from, and which signal wins
A search engine wants one canonical address for each thing. Sites, left alone, generate many addresses for the same thing, and every duplicate splits the ranking signal that should have pooled on one URL. The duplication comes from ordinary machinery: tracking parameters that create a new URL for every campaign, sort and filter parameters that reorder the same list, trailing-slash and uppercase variants the server treats as distinct, HTTP and HTTPS and www and non-www all resolving, session identifiers. Each hands Google several URLs with near-identical content, and Google has to guess which one you meant; when it guesses, it sometimes keeps the parameter-stamped version and buries the clean one.
Canonicalization is the set of signals that resolve the ambiguity: the canonical tag, redirects, internal linking, the sitemap. Google’s documentation lists the tag as one signal among several and calls it a hint, and in practice it is the weakest of the four, because it is the only one the crawler can disregard without consequence. Internal links are the strongest, because they are what the crawler follows. So the fix order for a canonical conflict is links first, sitemap second, tag last; a site that changes the tag and leaves ten thousand internal links pointing at the parameter version has changed nothing Google acts on. The contradiction I see most is a tag that says one URL while the internal links, the sitemap and the redirects vote another, so Google gets four instructions and follows none of them reliably.
I map the URL space from the crawl, flag the duplication patterns, and check Google’s verdict against the site’s in two places: the Duplicate, Google chose different canonical than user status, and the “Google-selected canonical” field in URL Inspection on a sample of each template. Two URLs from the same site ranking, badly, for the same term is how it surfaces in the results.
Crawl depth and orphan pages
Internal links do two jobs at once. They are the roads the crawler travels to discover pages, and they are votes that tell the engine which pages you consider important. A page with no internal links pointing at it is both undiscoverable and, in the engine’s eyes, unimportant.
Crawl depth is the number of clicks from the homepage to a page, and a rough proxy for how the engine weighs it: the deeper it sits, the less often it is crawled and the less authority flows to it. A proxy, because Google also enters through the sitemap and through external links, so a page four clicks deep with a strong external link is less buried than the number says. One correction I apply on commercial sites: the homepage is the wrong origin. On an ecommerce or multi-service site the entry points are the category and service pages, so I read depth from the top-level category as well as from the home, and a product that is six clicks from the homepage and two from its category is not the problem the site-wide histogram makes it look like. The product that is four clicks from its own category is.
The crawl gives the whole picture at once: the Crawl Depth distribution per template in the Site Structure tab, the filter Inlinks = 0 for orphans, and the redirect and 404 destinations for dead ends where a link points at a chain or a missing page. The output is the list of revenue URLs that are orphaned or buried, ranked by how much they matter.
Status codes, redirect chains and the XML sitemap
Status codes are the crawler’s first read on every URL, and the patterns matter more than any single code. A wall of 404s where content used to live means links and ranking history dead-ending into nothing. Redirect chains and loops: A to B to C to D. Gary Illyes said in 2016 that 3xx redirects no longer lose PageRank, so the cost of a chain sits elsewhere: Google’s documentation says Googlebot follows up to ten hops and then gives up, every hop is a fetch spent on nothing, and a chain built over three migrations usually has a loop or a 404 somewhere in the middle that nobody has walked.
WHAT A SITEMAP DOES
A URL in the XML sitemap tells Google the URL exists and you’d like it discovered, a suggestion for the discovery stage. Google decides indexing on its own. Of the sitemap’s fields, Google’s sitemap documentation says it reads lastmod if it is consistently accurate and ignores changefreq and priority. The hard limits are 50,000 URLs and 50 MB uncompressed per file; past that you need an index file. A sitemap full of URLs you’ve since deleted, redirected or noindexed feeds the discovery stage bad information and spends crawl on dead ends.
I audit a sitemap by asking whether every URL in it returns 200 and deserves to be there. A sitemap where a third of the URLs are redirects, 404s or noindexed pages points the crawler at things that waste its visit. The work is the full status-code map from the crawl, every redirect traced to its final destination to catch chains and loops, and the sitemap validated URL by URL against live status.
What structured data does for retrieval
Structured data is markup, usually JSON-LD, that states in machine-readable terms what a page is about: an article, by this author, on this date; a product, with this price and availability; an organization, and the entity behind it. It changes nothing the user sees and removes the engine’s guesswork about what the page means, which is what makes a page eligible for rich results and gives the answer engines a machine-readable entity to resolve. Everything above gets the page discovered, fetched, rendered and indexed; this layer decides whether the engine understands it, and understanding is what separates a page that is merely indexed from one the engine retrieves for the right query.
The audit reads three things. Is the schema present on the templates where it matters, and is it valid; invalid JSON-LD is simply not parsed, so a template with a broken block has no structured data at all. Two validators, and they disagree on purpose: Google’s Rich Results Test checks only the types Google uses and says whether a rich result is possible; the Schema.org validator checks syntax against the whole vocabulary, and markup that passes one and fails the other is itself a finding. Does the marked-up data match the visible page, since schema claiming a price or a rating the page doesn’t show is what Google’s spammy structured markup manual action is for. And is the entity information consistent: does the Organization, author and product markup tell one story across the site or three. Organization markup is often missing or on the homepage only, and author Person markup with a sameAs to a profile is rarer still; adding both is the cheapest fix in the whole audit.
Product and offer markup for shopping surfaces, and localized schema for multi-region sites, are covered in the international SEO audit and the ecommerce one; the general layer, valid and consistent markup that lets the engine understand any page, is here.
Core Web Vitals per template, and the crawl connection
Core Web Vitals are Google’s three measures of page experience, each with a threshold, plus a server-response metric that sits underneath all three.
| Metric | Good | Template that fails it most |
|---|---|---|
| Largest Contentful Paint | ≤ 2.5 s | Article and landing pages with a hero image |
| Interaction to Next Paint | ≤ 200 ms | Commercial templates with a tag manager |
| Cumulative Layout Shift | ≤ 0.1 | Listing pages with unsized images and injected banners |
| Time to First Byte | ≤ 800 ms | Any uncached dynamic template |
One detail changes how the field data reads. It comes from the Chrome UX Report, which only holds numbers for URLs with enough traffic; on a long-tail template the per-URL data is empty, PageSpeed Insights falls back to origin-level figures, and a failing template hides behind a passing homepage. So I read per template, and per template means sampling: twenty URLs from each class through the CrUX API, or the site’s own RUM. A single site-wide score is an average that hides which template is red. The INP failure on a commercial template is usually the tag manager and what it loaded, which is why the INP row above names it.
Speed is also a crawl concern, and that reading is the one most performance reports leave out. A slow server response directly limits how many pages Googlebot fetches per visit; the crawler budgets time as well as requests, and Google’s crawl-budget guidance says outright that response speed and server errors are what move the crawl limit up or down. The same guidance describes a lever most sites leave unpulled: answer a conditional request with 304 Not Modified when the page hasn’t changed, and the fetch costs almost nothing. That needs a usable Last-Modified or ETag on the template, which most CMS output doesn’t carry.
What the server logs settle
Server logs are the record of every request Googlebot made, every URL it fetched and every status it got back, in order. Everything above them in this article is inference from reading the site the way a crawler would; the logs replace the inference with the record, and they settle questions nothing else can. Which pages Googlebot crawls often and which it hasn’t touched in months. What share of its requests hit facet grids, parameter junk and redirect chains rather than the URLs that carry revenue. What share got a 304 back, which on most sites is zero, and that one number is the crawl-efficiency finding.
Two things come before the reading. Verify the Googlebot hits against Google’s published ranges, the googlebot.json file under developers.google.com or a reverse DNS to googlebot.com, because spoofed Googlebot traffic is common enough to distort the split. And separate Googlebot from the AI crawlers, which behave differently: Cloudflare Radar’s crawl-to-refer ratios for the 28 days to 21 July 2026 had Google fetching 4.6 pages for every visitor it sent back, OpenAI’s crawlers 217 and Anthropic’s 2,237. The ratios move month to month and the gap doesn’t, and on a site with a strict crawl allowance the AI crawlers are spending server capacity that Googlebot’s own limit responds to.
With those two done, the logs turn “your crawl budget is probably wasted” into a measurement: a share of verified Googlebot requests over a month hit URLs you don’t want indexed, and you can name them.
Retrievability: what the AI surfaces read
Retrievability is whether a passage from an indexed page can be pulled into an answer, and it became a separate check once answers started being assembled from passages. Google’s documentation says AI Overviews and AI Mode draw on the same index as Search, that there is no special markup for them, and that the Google-Extended robots token controls Gemini training and grounding and has no effect on inclusion in AI Overviews. So the question reduces to three technical ones: is the page indexable, is the passage in the raw HTML or only in the rendered DOM, and does it sit under a heading that says what it answers. That is why rendering and structured data sit at the center of the audit in 2026 and the old checklist of titles, meta tags and a clean robots.txt is what every platform now handles out of the box.
The standard checks, filed by failure state
Every standard check is still worth running; what changes is where each lands on the path, which decides whether a failure is a footnote or the reason a template is invisible. One row per check, where I read it, and the state it breaks.
| Check | How to detect | Failure state | What it means |
|---|---|---|---|
| HTTPS and mixed content | Crawl: Insecure Content report; DevTools Security panel | Uncrawlable / Misunderstood | HTTP pages still resolving, or HTTPS pages loading HTTP assets, split the URL space |
Meta robots and X-Robots-Tag | Crawl: Directives tab; response headers | Unindexable | A noindex in an HTTP header is invisible to anyone reading the page source |
| Self-referencing canonical | Crawl: Canonicals tab, Missing / Non-indexable canonical | Misunderstood | Absent canonicals leave the parameter variants to Google’s guess |
| Hreflang | Crawl: Hreflang tab, missing return links, non-canonical | Misunderstood | Reciprocal pairs, x-default, self-canonical on every copy |
| 404 vs 410 | Crawl: Response Codes; GSC Not found (404) | Undiscovered | John Mueller has said Google treats them nearly alike, with 410 dropping a little faster. The routing decision matters; the code follows |
| Pagination | Crawl: pages reachable only past page one; rel="next" unsupported since 2019 | Undiscovered | Only an <a href> to page two makes page two exist |
| Compression | Response headers: content-encoding: br or gzip | Inefficient | Uncompressed HTML on a large site is crawl budget spent on bytes |
| HTTP/2 or HTTP/3 | DevTools Network panel, Protocol column | Inefficient | Googlebot has crawled over HTTP/2 since November 2020; multiplexing lowers the cost of each fetch |
| DOM size | Lighthouse: Avoid an excessive DOM size (over ~1,500 nodes) | Unrenderable | Oversized DOMs slow rendering and push CLS, and they come from a template rather than a page |
| Mobile parity | Crawl with a mobile UA against desktop; compare content and links | Misunderstood | What the mobile version lacks, the index lacks |
| Lazy-loaded images and content | Crawl with rendering on; images missing src in raw HTML | Unrenderable | Native loading="lazy" renders; scroll-triggered loaders never fire for the crawler |
| Server errors | GSC Crawl stats by response, 5xx share; logs | Inefficient | Sustained 5xx is one of the two inputs Google names for lowering the crawl limit |
A site can pass all twelve and still have its best template sitting in Crawled – currently not indexed, and a site can fail three and rank, because the failures sat on pages nobody needed.
The report: issue, evidence, impact, fix, owner
A flat issue list sorted by severity is a crawler’s export, and its severity ranking is close to meaningless, because the crawler has no idea which URLs touch the business. It flags a critical warning on a page nobody reaches and a medium one on the template that carries the traffic, and hands you the dead page first. The report inverts that, and every finding carries five things.
- Issue. What’s wrong, in one line, tied to which of the six failure states it belongs to.
- Evidence. How it’s known: the crawl diff, the log segment, the index-check mismatch. The reading itself, reproducible.
- Impact. Severity re-weighted by which URLs it hits. A medium issue on the money template outranks a critical one on a page that was never going to rank.
- Fix. The correction, specific enough to ticket.
- Owner and recheck date. A robots rule is SEO’s, a rendering failure is engineering’s, a redirect chain belongs to whoever owns the server config. And a date to re-crawl and confirm the fix landed.
A report that skips the owner column doesn’t get acted on, because nobody knows whose job it is. One that skips the evidence column can’t be checked. The five together are what make it an engineering spec a dev team can act on Monday. A finished audit usually comes down to a handful of findings that matter, each on one of the six states, each with the evidence that proves it, with the rest of the crawler’s list in an appendix. If you want that read on your own site — the crawl, the render diff, the log segmentation, the index check — that is our SEO audit.
Frequently asked questions
How is a technical SEO audit different from an SEO audit?
The technical audit stops at whether the engine can reach, build, understand and retrieve the page; the full audit adds whether the page deserves to rank once it is there. Most sites need the technical read first, because a content fix on a template Google isn’t indexing changes nothing.
What gets rechecked between full audits?
Three things, on three cadences: the Pages report by template, monthly; a rendered crawl after any deploy that touches a template; and the log split, quarterly, on sites where crawl budget is a live constraint. A site that hasn’t changed structurally doesn’t need the full audit again. Its templates drift, and the three rechecks catch the drift.
How long does Google take to render a JavaScript page?
Google’s last public figure is from 2019: Martin Splitt, at Google I/O, put the median time between crawl and render at about five seconds, with a tail into minutes. It hasn’t published a newer one, and I don’t have a measurement of my own that separates queue time from crawl interval, so the answer in 2026 is that nobody outside Google knows. The audit treats any content that exists only after render as at risk, however fast the queue happens to be on a given day.