An enterprise SEO audit is a review of the templates, URL rules and crawl allowances that generate a large site’s pages. It reads the site per template, to find which rule produces a failure and which team owns the fix. The unit changes because the pages are output: one template renders every URL under it, so one wrong rule ships thousands of times the day it deploys, and one fix repairs them together.
The shape of the finding is always the same. Click through a hundred pages of a large catalog and every one of them looks fine: title present, canonical present, 200 status, indexable. Group the same URLs by the template that generated them and read the template as one object, and the fault is immediate. Every page under /catalog/ carries a canonical pointing at its parent category, and in Search Console the Duplicate, Google chose different canonical than user count on that one URL pattern starts climbing from the day the template deployed. Nobody touched a page. The deploy date set against the status curve is the evidence a page-by-page crawl can’t produce, because the crawl has no idea when anything happened.
Everything below is how I run that reading on sites in the low millions of URLs, which is the ceiling of what I work on; a hundred-million-page site is a different problem and I won’t pretend to cover it. It is also only the finding. Rebuilding the template, rewriting the URL rule or reconfiguring the redirect map is the build, and I mark where a finding turns into one.
What counts as an enterprise site
Enterprise is a combination of traits: many templates, many teams, several platforms stitched together, multiple markets, a CMS that fragmented over years, legacy layers nobody wants to touch. A 200,000-URL store run by one team on one clean CMS is easier to audit than a 40,000-URL site split across three content systems, four teams, and two acquisitions that were never merged.
Two traits define it. Propagation: one template edit lands on thousands of pages at once, for good or ill. Split ownership: the template belongs to engineering, the content to content ops, the redirect rules to the platform team, and no one person can see or fix the chain end to end.
| Dimension | Standard audit | Enterprise audit |
|---|---|---|
| Unit of analysis | The URL, the page | The template, the rule that stamps URLs out |
| A defect | One page to fix | One rule → N pages, fixed or broken together |
| Index coverage | One site-wide number | Broken out per template, per sitemap |
| Crawl budget | One site, one allowance | One allowance per site, shared across every Google crawler |
| Evidence | The crawler’s flag | Deploy date against the Search Console curve; logs against the crawl |
| Deliverable | Flat list of URL-level issues | Causes rolled up, weighted by revenue, owner named |
One correction before the thesis hardens into a rule. Not every URL on an enterprise site is template output. The homepage, the top categories, the flagship guides and the handful of pages that carry a disproportionate share of revenue don’t obey any template, and they get read page by page, the old way. Knowing which few hundred URLs to exempt is part of the audit.
Tools by scale
The crawler and the log reader both change at roughly a million URLs; Search Console changes with the number of hosts and folders.
| Job | Into the low millions of URLs | Past that |
|---|---|---|
| Crawl | Screaming Frog in database storage mode, full crawl, JavaScript rendering on for the templates that need it | Lumar, Botify or OnCrawl; or a per-template sample validated against logs and Search Console |
| Logs | Screaming Frog Log File Analyser, into the millions of lines | Botify, or a warehouse query (BigQuery, Snowflake) over the raw access logs |
| Search Console | A URL-prefix property per host for Crawl stats; a property per major template folder for Pages and Performance | The same, plus the Bulk data export to BigQuery for query-level data past the 1,000-row cap |
| Index status | URL Inspection API, sampled per template (2,000 requests a day per property) | Per-template sitemaps and their submitted-versus-indexed ratio, read continuously |
My position on crawl scope: into the low millions a full crawl is still practical and worth running, because sampling can’t find an orphaned template. Past that, a desktop crawler stops being the tool, and the alternative is a rotation of per-template samples. A pattern shows up in a sample; the URL class you didn’t know existed doesn’t, which is why the log segmentation below runs against every host regardless of how the crawl was scoped.
Template-level analysis: crawl segmentation by URL pattern
You could optimize a thousand pages by hand and still not touch the template breaking a hundred thousand, so the audit reads the templates and rules that stamp the pages out. I crawl, then segment every URL by template (usually by URL pattern, sometimes by a CSS selector or a data-layer attribute when the pattern doesn’t reveal the template) and read each template’s SEO fingerprint as one thing: canonical behavior, index status, title pattern, meta robots, internal links in and out, all rolled up to the rule that produced them. A healthy template has one fingerprint. A broken one has one fingerprint too, which is why the break is visible.
The failure I find most often is the canonical one from the opening: a category template that points every page at a parent, or at itself with a stripped parameter that collapses distinct pages into one, and a directory falls out of the index. Second is the header-level noindex, an X-Robots-Tag: noindex set on a whole directory in the server config, invisible to anyone reading page source, which is how a staging rule survives into production. Once the template is named, the finding is complete; rebuilding it is engineering’s ticket.
Search Console at enterprise scale
Search Console is built for a site, and an enterprise is several. One domain property and one Performance report average most of the findings above out of existence, and the limits below decide the setup.
| Limit | Value | What it forces |
|---|---|---|
| Performance UI rows | 1,000 per view | Bulk data export to BigQuery (available since 2023) for query-level data on a site with a million queries |
| Retention | 16 months | A two-year baseline needs the export to have been running already |
| URL Inspection API | 2,000 requests a day per property | Live index status is sampled per template, never read URL by URL |
| Crawl stats report | Domain properties and host-root properties only | Crawl split by host comes from a property per host; crawl split by template comes from the logs, not from Search Console |
So the setup is a URL-prefix property per host alongside the domain property, which gives Crawl stats per host, and a property per major template folder, which gives the Pages report and Performance per template. The two exclusion statuses that matter most read differently: Discovered – currently not indexed clustering in one template is the crawl-starvation signal, while Crawled – currently not indexed is, in practice, a quality or duplication verdict. Most reports lump them, and they point at different owners.
Crawl budget: server log analysis by host and template
Google’s crawl-budget guidance gives each site its own allowance, so www.example.com, shop.example.com and m.example.com are read as separate crawl-budget problems, and in my audits the smallest host is the one that starves. Since the 22 July 2026 rewrite of the guidance, the capacity limit is shared across all of Google’s crawlers — Googlebot, AdsBot, the image and Shopping fetchers — and every site starts on a conservative default that grows only as the server proves it can take more. My reading of that for an enterprise running heavy Shopping feeds and ads verification: the ads crawler is spending the organic crawler’s budget.
The pattern in the logs
“Millions of URLs” stays an abstraction until you pull the server logs and watch where Googlebot spends its day. A large share of what verified Googlebot touches returns nothing worth indexing, and the most-crawled URLs are dominated by filter and sort combinations that aren’t indexed and have no search volume, hit more often than the category pages they hang off. The parameter strings name the platform:
| Platform | Facet or variant string in the logs |
|---|---|
| Salesforce Commerce Cloud | ?prefn1=refinementColor&prefv1=Blue |
| Magento layered navigation | ?color=53&price=50-100 |
| Shopify Plus | filter.v.option.color=blue |
| Sitecore | ?sc_lang= variants |
| Adobe Experience Manager | .html selector chains |
What I don’t have is a clean benchmark for that waste share. Published log analyses describe individual sites, and the careful ones say so; a population number doesn’t exist, and I’d distrust anyone who quotes one.
Verification, AI crawlers, segmentation
Before trusting the logs I verify the Googlebot hits against Google’s published IP ranges, the googlebot.json file or a reverse DNS to googlebot.com, because spoofed Googlebot traffic on a large site is common enough to distort the split. I separate Googlebot from the AI crawlers (GPTBot, ClaudeBot, PerplexityBot and the rest), which behave differently and on an enterprise origin can outnumber it. Verified hits then get segmented by template, the same segmentation as the crawl, so the two datasets line up: where the crawl says the templates are, and where the logs say the bot went. The output is a share of Googlebot’s month spent on URLs you don’t want indexed, and the names of the templates that took it.
Index bloat, robots.txt and XML sitemaps by template
The enterprise index question is what share of indexed URLs is worth indexing. Programmatic generation is the usual cause, because generation and value are different things: a template that spins up a page for every city, every filter combination, every attribute pair can manufacture tens of thousands of URLs that each add nothing the index didn’t already hold. Since March 2024 that has a name in Google’s spam policies, scaled content abuse, generating many pages primarily to manipulate rankings regardless of how they were produced, and the August 2025 spam update was widely read as the enforcement pass against thin, near-duplicate, programmatic sets. A template-generated tail of location or attribute pages is the textbook case.
I group indexed URLs by template and ask, template by template, which ones earn their place and which inflate the count. Robots.txt and sitemaps fold in because they are where the fix gets aimed. Robots.txt at scale is a decision about which whole templates to keep bots out of, with one reminder: Google reads the file up to 500 KB, and an enterprise robots.txt that grew by accretion can pass that and have its tail silently ignored. XML sitemaps get segmented by template, one sitemap per template under the 50,000-URL and 50 MB limits, gathered by a sitemap index. When a sitemap holds one template’s URLs, Search Console’s indexed-versus-submitted ratio per sitemap says which template is failing to index, without a crawl. Setting that up costs a developer an afternoon, and it is the cheapest ongoing monitor an enterprise site can run.
Striking-distance keywords and cannibalization by template
Striking-distance keywords are the ones sitting at roughly positions 5 through 20, filterable out of Search Console, and every rank tracker has the report. What the report doesn’t do is the second half: telling you which of those pages can move and which only look close. A page at position 30 needs a rebuild. A page at position 8 with demand behind it needs a nudge: a few internal links from high-authority templates, a couple of targeted external links, one metadata and content pass to cover the entity properly. I surface candidates by crossing rank data against the demand map with a floor on search volume, then read each one for intent match, inbound links that carry weight, and whether the content covers the entity or only mentions the keyword. The cap on a near-win is usually the template itself (a canonical rule or an internal-linking pattern that limits every page it generates), which sends the finding back to the template section.
Cannibalization at this scale is a template problem too. When two templates generate pages for the same intent (a category and a filter page, a location page and a service page) the ranking URL alternates between them and neither holds a position. I read it in the Performance report by query with the Pages dimension: a query whose top URL changes month to month is being served by two templates, and the fix is a rule, a canonical or a noindex on the template that shouldn’t compete, applied once.
Old templates and entity coverage
Enterprise sites are old, and old sites carry old optimization fossilized into their templates: the target term in the title, the H1, and three or four times in the body, stamped across thousands of URLs by a template built when keyword density was the accepted play. Google has described its ranking as reading pages in context since BERT in 2019, and a page built around how many times it says the term underperforms a page that covers the field the query implies. The check per template is entity coverage against what ranks, which entities the winning pages co-mention that the template never does, and an information-gain read of whether the template adds anything the ranking set doesn’t already hold. Bi-encoder and late-interaction models are how I reason about that similarity; they are a model of relevance, not a description of Google.
Site architecture: taxonomy, navigation and the internal link graph
Large organizations build the site to mirror themselves: business units, product lines, the org chart, the way the company files its own thinking. That is rarely how the market searches, so demand lands in the wrong place or finds no place at all. The fastest check is the Search Console query report set against the navigation labels: if the words bringing impressions don’t appear in the menu, the demand and the pages exist and never meet. Then the internal link graph at scale, read per template from the crawl’s Site Structure tab and the Inlinks = 0 filter: commercial templates orphaned because no other template links to them, whole clusters isolated from the rest of the site, authority pooling in the blog, which everyone links to by habit, instead of the money templates that need it. Reading the graph is a tooling job; restructuring the navigation around demand is a business decision the audit informs.
Duplicate content, scaled content and site reputation abuse
Systemic duplication
Duplication on an enterprise site is a template producing near-identical pages by the thousand: location pages that differ only by the city name dropped into the same boilerplate, filter pages, product descriptions pulled verbatim from a manufacturer feed across every SKU. Google’s helpful-content system has been a site-wide classifier since it launched in August 2022, and the March 2024 core update folded it into core ranking, so it no longer runs as a separate update you can see coming. A bloated tail of thin, templated pages drags down the domain, including the templates that sell, and recovery lands at a later core update; the industry’s read of the sites hit in 2023 is that recoveries mostly took a year or more. I score the thin templates against the SERP they’d have to compete in: does this entire template, at its full scale, add anything the ranking results don’t already have. When the answer is no, the template is a liability.
Third-party sections
There is a third-party version of this that is specific to enterprise, and it comes straight from split ownership. Since May 2024 Google has enforced a site reputation abuse policy against third-party content hosted on a site to exploit its ranking signals — coupon sections, review sections, partner content on a subdomain or in a folder — and the November 2024 clarification closed the loophole: first-party oversight of that content doesn’t exempt it. On 15 May 2026 the spam policies were rewritten so that “attempting to manipulate generative AI responses in Google Search” counts as spam too, with the June 2026 spam update the first rollout under the new wording. Most enterprise sites I’ve audited had at least one of these sections, usually inherited from an acquisition or a partnership nobody in SEO was told about. The audit inventories them with an owner per section, because the team that owns them is rarely the team that will be answering for them.
Technical SEO pass: Core Web Vitals, JavaScript rendering, canonicals, schema
The standard technical checks get one tight pass here, because they rarely surprise and shouldn’t pad the audit. Core Web Vitals read per template, never as one site-wide score: a PDP template and a PLP template fail differently, and the Chrome UX Report only holds per-URL numbers for URLs with enough traffic, so on enterprise the long tail inherits the origin’s figures and a failing template can hide behind a passing one. JavaScript rendering asks whether Googlebot receives the same commercially relevant DOM the user does, and it matters more here than anywhere because these sites run the heaviest frameworks. Canonical hygiene across the URL space. Schema, and specifically whether Organization markup exists once, on every host, or three contradictory times. Mobile parity. The full reading of that layer — discovery, crawl, render, index as one path — is the technical SEO audit; this article assumes it.
Outside this article on purpose: the seams between locales on a multi-market site, covered in the international SEO audit; link building, since link equity gets a single check here (which templates the external links land on); and the platform migration itself, the most common trigger for an enterprise audit, which is a build project with its own checklist. The audit reads what the migration left behind.
Governance and audit cadence: owner per template
On enterprise, “what’s broken” is half the answer, and “who controls the system that’s broken” is the other half. A fix with no named owner dies in an unworked backlog no matter how correct the diagnosis, so every finding names the system and the team that runs it.
Cadence follows from that. A full audit annually, which on a site with logs and the bulk export already in place takes four to eight weeks, and the spread is explained by the number of templates, not the number of URLs. A template-level health check after every deploy to a high-traffic template. Continuous monitoring in between, where the per-template sitemap ratio from the index section is the alarm: when a template’s ratio drops after a deploy, whoever owns that template knows within days. Unscheduled triggers are a platform migration, a design-system change, a CDN or hosting move, an acquisition folded into the domain, and an unexplained double-digit drop in non-brand organic sessions.
Enterprise SEO audit report: causes, revenue weight, owner
A flat 500-issue export is worse than useless at this scale, and the reason is structural. It lists URL-level symptoms of template-level causes — tens of thousands of rows that are, underneath, one wrong canonical rule — and sorts them by a crawler’s idea of severity, which has no notion of which pages touch revenue.
The deliverable inverts that. It rolls symptoms up to their cause, so one rule and its N affected URLs collapse into one fix. It weights by business impact, because a medium flag on the template that drives revenue outranks a critical one on a template nobody should reach. It names the owning team per fix. And it carries a date for each, when the template deployed and when the index moved, so the fix can be verified against the same signal that exposed it.
Enterprise SEO audit checklist: checks by template
Where I read each layer and what surfaces the failure, with the template as the row key throughout.
| Check | Where I read it | What surfaces the failure |
|---|---|---|
| Template fingerprint | Full crawl (or per-template sample) segmented by URL pattern | One template, one wrong canonical or robots signal on every URL |
| Deploy evidence | Release log against GSC Pages report per property | Duplicate, Google chose different canonical than user climbing from the deploy date |
| Header-level noindex | Response headers per template; crawl Directives tab | X-Robots-Tag: noindex on a directory nobody meant to hide |
| Crawl allowance | Verified Googlebot hits in logs, by host and template | Facet and sort URLs out-crawling the categories they hang off |
| Crawl starvation | GSC Crawl stats per host; Pages report per template folder | Discovered – currently not indexed piling up in one class |
| Index bloat | Indexed URLs grouped by template; per-template sitemap ratio | A generated tail indexed at scale with no demand behind it |
| Striking distance and cannibalization | GSC positions 5–20 crossed with the demand map; queries whose top URL alternates | Near-wins capped by a template rule; two templates serving one intent |
| Entity coverage | Co-mentioned entities and information gain per template vs the ranking set | Keyword-density templates covering none of the field |
| Taxonomy vs demand | GSC query report against navigation labels; Site Structure and Inlinks = 0 per template | Impressions on words the menu never uses; orphaned commercial templates |
| Third-party sections | Inventory of partner folders and subdomains; ownership per section | A coupon or review section carrying the domain’s authority and nobody’s name |
If you’d rather have that read on your site than run it yourself, that is our SEO audit.