All SEO Articles
Technical SEO

The SEO Zombie Pages Eating Your Crawl Budget

Thousands of low-value URLs may be draining Google’s attention while your most important pages wait to be discovered. Learn how to spot, contain, and eliminate SEO zombie pages.

11 minute read
a wooden block that says seo on it
Photo by NisonCo PR and SEO on Unsplash

Your website may be hosting a zombie outbreak.

Not the fun kind with leather jackets and poor life choices. The expensive kind: thousands of URLs shuffling through your server logs, mumbling “200 OK,” eating crawl resources, and contributing absolutely nothing to rankings, leads, or revenue.

These are SEO zombie pages: URLs that technically exist, often return a valid status code, and may even be discoverable by Google—but have no business being crawled regularly, indexed at all, or included in your growth plans.

For founders, this is less “spooky SEO folklore” and more “why is Google spending time on our /search?q=blue+widgets&page=47 page while our new product launch still isn’t indexed?”

Let’s grab a flashlight and inspect the URL graveyard.

What makes a page an SEO zombie?

A zombie page is not simply a page with low traffic. Plenty of valuable pages begin life unnoticed. A new pricing comparison page, a niche integration page, or an excellent help article may attract only a handful of visits before it finds its audience.

A zombie page is a URL that consumes crawl attention or pollutes index signals without delivering meaningful search value.

Common traits include:

  • Little or no organic traffic over a meaningful period
  • No backlinks, conversions, or user engagement worth preserving
  • Thin, duplicated, or near-duplicated content
  • Endless variations of the same page
  • URLs generated by filters, internal search, pagination, tags, or tracking parameters
  • Pages with no clear search intent
  • Pages that should be consolidated, redirected, blocked from crawling, or marked noindex

Think of your website as a restaurant kitchen. Your best pages are the dishes customers order. Zombie pages are 400 bowls of garnish being prepared every night because someone accidentally put “parsley variants” on the menu.

Googlebot is not emotionally attached to your parsley.

Crawl budget: important, but not a reason to panic

“Crawl budget” is the amount of crawling Google is willing and able to devote to your site over time. It is shaped by factors such as:

  • Your site’s size and URL volume
  • Internal linking and URL discovery patterns
  • Server performance and crawl capacity
  • Perceived demand for your content
  • How many low-value URLs Google encounters

For a small, clean site with a few hundred pages, crawl budget is rarely the main bottleneck. If Google is not indexing a great new page, the cause might be weak internal linking, poor content quality, duplication, or simply insufficient time.

But crawl waste becomes very real when a site has tens of thousands—or millions—of URLs, especially in ecommerce, marketplaces, SaaS platforms, publishers, directories, and sites with enthusiastic filter systems.

The issue is not that Google has a tiny ration card and you stole its last bread coupon. The issue is that you are making its crawler walk through a maze of doors that all lead to the same broom cupboard.

a laptop computer sitting on top of a desk
Photo by Lukas Müller on Unsplash

The usual suspects: where zombie URLs come from

Zombie pages rarely rise from the earth on their own. Your CMS, framework, product team, or “just one quick filter” feature usually gives them a hand.

Faceted navigation and filter combinations

Faceted navigation is enormously useful for users. It is also capable of generating enough URLs to make an enterprise crawler reconsider its career choices.

Imagine an online store selling office chairs. Users can filter by:

  • Color
  • Material
  • Brand
  • Price
  • Ergonomic features
  • Delivery speed
  • Rating
  • Availability

Now combine them:

/chairs?color=black&material=mesh
/chairs?color=black&material=mesh&rating=4
/chairs?color=black&material=mesh&rating=4&delivery=next-day

Some combinations might serve real search demand. “Black mesh office chairs” is plausible. “Black mesh office chairs rated 4+ with next-day delivery under $300” is usually a page nobody searches for, including the customer who created it thirty seconds ago.

The danger grows when every filter combination is internally linked, crawlable, indexable, and included in XML sitemaps. That is not navigation. That is URL confetti fired directly into Google’s face.

Internal site-search results

Internal search pages can be useful for visitors, but they are often poor landing pages from Google.

Examples:

/search?q=crm
/search?q=best+crm
/search?q=crm%20
/search?q=CRM&utm_source=email

Search-result pages can create duplicate results, thin content, nonsensical queries, and occasionally embarrassing URLs generated by users. Google has long advised against allowing internal search result pages to be indexed.

Usually, the sensible approach is:

  • Keep them available to users
  • Apply noindex, follow where appropriate
  • Avoid linking them broadly in crawlable ways
  • Ensure parameter variations do not create infinite duplicates

Do not block them in robots.txt instead of using noindex if they are already indexed. If Google cannot crawl the page, it may not see the noindex directive. First allow discovery of the directive; then manage crawling deliberately once URLs have dropped from the index.

Tag, category, and archive pages with no job

Tags are where content strategies go to acquire extra limbs.

A blog may have useful topic hubs like /tag/technical-seo/. Fine. But it may also have:

/tag/google/
/tag/seo-tip/
/tag/seo-tips/
/tag/search-engine/
/tag/search-engines/
/tag/2023/

If these archives contain only one or two posts, duplicate each other, or have no distinct purpose, they create weak pages at scale.

The same applies to empty categories, expired collections, outdated event archives, author pages with no unique information, and paginated archives that merely repeat snippets without helping users find anything.

Parameters, sorting, session IDs, and calendar traps

Other prolific undead URL factories include:

  • ?sort=price-asc and other sort orders
  • Tracking parameters such as utm_source
  • Session IDs appended to URLs
  • Printer-friendly versions
  • Alternate formats
  • Calendar pages with “next month” links extending toward the heat death of the universe
  • Infinite scroll implementations that create crawlable state URLs
  • Case variations, trailing slash variants, and HTTP/HTTPS duplicates

Each URL may look harmless. Together, they become a family reunion of pages whose only common interest is consuming server resources.

a scrabbled wooden block with the word stem on it
Photo by NisonCo PR and SEO on Unsplash

How to find the undead without deleting your best pages

Do not start by exporting every URL with fewer than ten visits and swinging an axe. That is how you accidentally redirect your most valuable niche comparison page to the homepage, which is the SEO equivalent of treating a paper cut with a flamethrower.

You need to combine several views of the site.

Start with Google Search Console

Use the Page indexing report to spot patterns such as:

  • Crawled – currently not indexed
  • Discovered – currently not indexed
  • Duplicate without user-selected canonical
  • Alternate page with proper canonical tag
  • Excluded by noindex
  • Blocked by robots.txt
  • Soft 404

None of these labels is automatically a problem. “Alternate page with proper canonical,” for example, can be exactly what you intended. The value lies in the volume, URL patterns, and whether the outcome matches your strategy.

Look for repeated path structures and parameter patterns:

/filter/
/search?
?sort=
?color=
/tag/
/page/

Then ask: “Would we proudly show this URL to a potential customer arriving from Google?” If the answer is an uncomfortable pause followed by “well, technically…,” you have a lead.

Crawl your site like a bot would

A crawler such as Screaming Frog, Sitebulb, or a comparable platform can reveal:

  • Indexable URLs with low word counts
  • Duplicate titles and meta descriptions
  • Duplicate or near-duplicate content
  • Canonicalised URLs still linked internally
  • Orphan pages, if you combine crawl data with analytics and sitemap data
  • Paginated and parameterized URL patterns
  • Broken internal links and redirect chains

Export URL data and group it by directory, template, or parameter. A list of 80,000 URLs is terrifying. A pivot table showing that 61,000 come from ?sort= is a plan.

Use server logs when crawl waste is serious

For larger sites, server log analysis is the closest thing to a security camera for Googlebot.

Logs tell you what search bots actually request—not merely what your tools think they might request. Look for:

  • High crawl frequency on parameter URLs
  • Repeated bot visits to redirected URLs
  • Crawl activity on non-canonical pages
  • 404 or 5xx responses being hit repeatedly
  • Important URLs receiving little or no bot attention
  • Crawl spikes following new navigation, filters, or releases

If Googlebot spends 40% of its requests crawling search pages and product sort orders, while newly launched product pages wait weeks for a crawl, you have evidence, not vibes.

Score pages by purpose, not vanity metrics

For each suspicious URL group, classify it:

QuestionWhat it helps determine
Does this page satisfy a distinct search intent?Keep or improve it
Is the content substantially unique?Canonicalize, merge, or remove duplicates
Does it receive meaningful organic traffic, links, or conversions?Preserve valuable exceptions
Is it needed for users but not search?Consider noindex
Is it an accidental URL variation?Redirect or canonicalize
Does it create an unlimited URL space?Restrict crawl paths and generation rules

The unit of analysis is usually not one page. It is a URL class: all filtered category pages, all internal searches, all sort variations, all empty tags.

That is where founders get leverage. You do not fix 50,000 URLs individually. You fix the machine that keeps giving birth to them.

Choose the right weapon: redirect, canonical, noindex, or robots.txt?

Technical SEO has several tools, and they are not interchangeable. Using them carelessly is like trying to repair a watch with four different hammers.

Use redirects when a page has truly moved or should no longer exist

A permanent redirect is best when there is a clear replacement.

Examples:

  • An old product URL replaced by the current equivalent
  • A merged article redirected to the stronger consolidated guide
  • Trailing slash or HTTP variants redirected to the canonical version
  • Expired campaign pages with an obvious relevant successor

Avoid redirecting every deleted URL to the homepage. That often creates a poor user experience and may be treated as a soft 404 by search engines. Redirect to the closest meaningful alternative—or return a proper 404/410 if none exists.

Use canonical tags for close duplicates you need to keep accessible

A canonical tells search engines which version you prefer for indexing when multiple similar URLs remain useful.

For example:

/chairs?color=black
/chairs?color=black&utm_source=newsletter

The tracked version can canonicalize to the clean URL.

Canonicals work best when the pages are substantially similar. They are hints, not remote-control commands. If your canonicalized pages contain meaningfully different products, text, or intent, Google may ignore your preference.

Also: do not canonicalize a page to an unrelated category because you wish the problem would go away. Search engines have seen this trick. They were not impressed.

Use noindex for pages users need but searchers do not

noindex is often ideal for:

  • Internal search results
  • Account and login pages
  • Thin tag archives
  • Filter combinations with no search value
  • Thank-you pages
  • Certain paginated or utility pages, depending on site architecture

A noindex page can still pass internal link equity if it remains crawlable and links are followed. But if you have millions of low-value URLs, allowing unrestricted crawling forever may still be wasteful.

Use robots.txt to prevent crawling—not to remove pages from search

robots.txt controls crawling. It does not reliably remove a URL from the index.

Use it carefully for URL spaces that should not be crawled, particularly when they create huge combinations:

Disallow: /search
Disallow: /*?sort=
Disallow: /*?sessionid=

Test rules thoroughly. An overenthusiastic Disallow can block JavaScript, CSS, product pages, or other things you would quite like Google to see.

A common sequence for a large unwanted URL set is:

  1. Stop generating or internally linking to the URLs.
  2. Add noindex where indexing must be reversed.
  3. Clean up canonicals and redirects.
  4. Once search engines understand the intended status, use crawl controls for genuinely useless URL patterns where appropriate.
  5. Remove those URLs from XML sitemaps.

Stop feeding the zombies: prevention beats cleanup

The cheapest zombie page is the one your platform never creates.

Put rules into product and engineering decisions

Before launching a new filter, archive, user profile, or programmatic landing-page template, ask:

  • Will this create one URL or one million?
  • Which variations deserve indexing?
  • Are default and sorted versions canonically controlled?
  • Can users create arbitrary indexable pages?
  • Will this feature introduce crawlable infinite paths?
  • Is the page useful without a logged-in session?
  • Does it have unique content beyond a database query?

These are not “SEO asks” to tack onto the final sprint. They are architecture questions. If a product manager can add 300,000 indexable URLs with a checkbox, SEO needs a seat in that meeting.

Keep XML sitemaps painfully honest

Your XML sitemap is not a wish list. It should contain canonical URLs you genuinely want crawled and indexed.

Do not include:

  • Redirects
  • 404s
  • noindex URLs
  • Canonicalized duplicates
  • Parameter variants
  • Thin filtered pages you would not defend in public

A clean sitemap helps search engines prioritize the pages you actually care about. More importantly, it gives you a useful diagnostic baseline: if a sitemap URL is not indexed, investigate; if an excluded URL appears nowhere in your sitemap, that may be completely fine.

Build internal links around priority pages

Google discovers and evaluates pages partly through links. If your strongest navigation points bots toward filters, tags, and utility URLs while your key commercial pages are buried five clicks deep, your site has arranged its own scavenger hunt.

Make sure important pages are:

  • Linked from relevant hubs and categories
  • Included in logical navigation
  • Supported by contextual internal links
  • Not dependent on a site-search query for discovery
  • Able to return a fast, stable 200 response

The goal is not to force crawlers to obey. The goal is to make the good paths obvious and the bad paths boring.

The prescription

  1. Export excluded and indexed URL patterns from Google Search Console. Group them by template or parameter.
  2. Crawl your site and identify thin, duplicate, canonicalized, redirected, and parameterized URL clusters.
  3. Check server logs if you have a large site or clear crawl delays; find where Googlebot actually spends time.
  4. Decide each URL class’s job: index, canonicalize, redirect, noindex, or prevent crawling.
  5. Remove zombie URLs from internal links and XML sitemaps.
  6. Fix the generator, not just the generated pages—especially filters, internal search, tags, and sort parameters.
  7. Monitor crawl activity and index coverage after release. Zombie outbreaks have a habit of returning through the next “tiny” product feature.