Skip to content
SEO

Ecommerce SEO: The Category, Filter and Duplicate Content Mistakes That Quietly Cost You Rankings

Picture a store with two thousand products and more than a hundred thousand URLs known to Google. Filters, sort orders and pagination created most of them. Googlebot spends its time on variations of the same category page and finds new products late. We went through Google’s documentation, tested the default setup of a large Central European ecommerce platform, and wrote down how to handle categories, filters and duplicate content so the right pages are the ones that rank.

Ecommerce SEO: about 50% of crawl issues reported to Google involve filters, one category with five filters and sorting produces 644,204 URLs, 73.56% of people in the EU shopped online in 2025

Online shopping is the norm across Europe. According to Eurostat, 73.56% of people aged 16 to 74 in the EU bought something online in 2025, and in the Netherlands and Ireland it was well over 90%. For online stores, that means crowded search results, where technical decisions matter as much as content: how many URLs a store generates, and which of them search engines are allowed to crawl.

People who shopped online in the last 12 months, 2025 (%)
  • Ireland95.05%
  • Netherlands94.18%
  • Czechia83.67%
  • Germany81.38%
  • Slovakia78.76%
  • EU average73.56%
  • Austria72.71%
  • Poland69.73%
  • Bulgaria51.20%

Share of people aged 16 to 74 who bought goods or services online at least once in the previous 12 months. Source: Eurostat, isoc_ec_ib20, updated April 17, 2026.

Here we focus on what sets an online store apart from a regular website: category pages, faceted navigation, sorting, pagination, product variants and duplicate content. For the broader picture of what works in search this year, see SEO strategy in 2026: what actually works.

01Category pages are your strongest landing pages

For a broad query like “men’s running shoes,” search results rarely show a single product page. They show listings. Build your categories around how people search. How your warehouse is organized doesn’t matter to them. A supplier may sort stock into “Footwear M” and “Footwear W,” but shoppers still type “men’s shoes” and “women’s shoes.”

Google doesn’t work out which pages matter from your URL structure. It looks at links. According to its documentation, the more internal links point to a page, the more important that page is relative to the rest of the site. Google recommends a clear hierarchy of menu, categories, subcategories and products, and notes that Googlebot doesn’t type queries into your search box. A product that no category links to may never be found.

  • Title and H1: the main query first, without repeating the store name in the heading.
  • A short intro above the product grid: two or three sentences that help people choose. Long blocks of text below the pagination rarely get read.
  • Links to subcategories: visible above the products instead of tucked away in a dropdown.
  • Every category earns its place: avoid two categories that show almost the same products.

02Filters create more URLs than you have products

Take a category with five filters (brand, color, size, material and price), each with ten values. If a shopper can pick one value or none from each filter, that is 11 × 11 × 11 × 11 × 11 = 161,051 combinations. Add four sort orders and you get 644,204 URLs. That is before pagination, before selecting several values of the same filter, and before parameters show up in a different order. For a single category.

Google calls this faceted navigation, and its documentation warns that the most common implementation, based on URL parameters, can generate infinite URL spaces. Crawlers waste time on useless pages and have less left for new products. Back in December 2024, Gary Illyes of Google wrote that faceted navigation is by far the most common source of overcrawl issues site owners report. In February 2026, on the Search Off the Record podcast, he put rough numbers on it:

Crawl issues reported to Google in 2025 (%)
  • Faceted navigation (filters)50%
  • Action parameters (add to cart, wishlist)25%
  • Irrelevant parameters (UTM, session IDs)10%
  • Plugins and calendars5%
  • Other URL oddities2%

Approximate shares of the reports site owners submitted through Google’s crawl issue form, not of all websites. Gary Illyes did not break down the remaining 8%. Source: Search Off the Record transcript, February 3, 2026.

About a quarter of the reports involve action parameters. The classic case is an “Add to cart” or “Add to wishlist” link that points to a URL with a parameter instead of using a button. Every such link on a product page doubles the number of URLs a crawler can find. Those URLs never belong in the index.

None of this means filters are a mistake. Shoppers need them, and even large stores often have too few. Baymard Institute benchmarked 343 top-grossing ecommerce sites in the US and Europe:

Where large stores’ product lists fall short (% of sites)
  • Not all 4 essential sort types (mobile)69%
  • Not all 4 essential sort types (desktop)68%
  • Not all 5 essential filter types51%
  • No overview of applied filters20%
  • Can’t combine values of one filter14%

Baymard’s essential filters: price, user rating, color, size and brand. Benchmark of 343 large US and European ecommerce sites, article updated September 9, 2025. Small stores are not included. Source: Baymard Institute.

So keep your filters for shoppers, and let search engines see only the combinations that deserve to rank.

03Which filter combinations to let into the index

Some filtered listings deserve their own landing page. People search for “Nike men’s running shoes” or “women’s winter jackets XL,” and the general category doesn’t answer those queries precisely. Most combinations, though, have no search demand and look like the parent category with a slightly shorter list.

Five questions before indexing a filter page: does anyone search for it, does the listing have enough products, can it have its own heading and intro, does the URL have a fixed format and filter order, does the site link to it
If the answer to any question is no, keep that combination out of the index.

For combinations you do want indexed, Google recommends using the standard & separator between parameters, keeping filters in the same order in the URL, and returning a 404 status code when a combination has no results. The same goes for URLs with duplicate filters, nonsensical combinations and pagination pages that do not exist. Google also advises against redirecting empty combinations to a generic “not found” page. The landing page itself then gets its own heading, intro text, a self-referencing canonical, a place in the sitemap and a link from the parent category.

Keep crawlers out of every other combination. Google’s first recommendation is to disallow crawling of filter URLs in robots.txt, or to put filters in the URL fragment after #, which search engines generally ignore when crawling. Canonical tags and rel="nofollow" are, in Google’s words, generally less effective in the long term. Nofollow also only works if every link to that URL carries it, including links from other websites.

Some platforms generate these landing pages for you. Shoptet, a popular Czech ecommerce platform, offers an Advanced SEO add-on that creates parametric categories from combinations of up to four parameters and preselects every combination that contains at least one product. That is where you need to cut. A combination with one product and no search demand adds a thin page to your sitemap, not new traffic.

04Sorting, pagination and the “Load more” button

Sorting by price or popularity shows the same products in a different order. It doesn’t belong in the index. Google’s pagination guide allows either a noindex tag or a robots.txt rule. For large catalogs, however, its crawl budget documentation recommends robots.txt, because a crawler has to fetch a noindexed page to see the tag, which wastes crawl capacity. Google says this mainly concerns sites with over a million pages that change weekly, or over 10,000 pages that change daily, and calls those numbers rough estimates. For a smaller store, noindex is usually fine.

Pagination is different. Page two and beyond list different products, so according to Google each page needs:

  • its own URL, such as ?page=2 or /page/2/, never just a fragment after #,
  • its own canonical URL, not one pointing to page one,
  • a plain <a href> link to the next page. Googlebot doesn’t click a “Load more” button.

Google no longer uses rel="next" and rel="prev". They do no harm and other search engines may still read them, but they will not fix anything in Google.

Watch out for pages that should not exist. On September 30, 2026, we opened page 99 (/dekorace/strana-99/) of the Dekorace (Decorations) category on Shoptet’s official demo store (classic.shoptet.cz). The category has eight products. The page returned a 200 status, an index,follow robots meta tag and an empty product list. Google recommends a 404 for nonexistent pagination. If nothing links to such a URL, a crawler won’t find it and nothing happens. Still, check your own store, since settings vary.

05Where duplicate content in an online store comes from

First, the reassuring part: there is no penalty for ordinary duplicate content in an online store. Google said so in 2008, and its current documentation states that some duplicate content on a site is normal and not a violation of its spam policies. Google groups duplicate URLs and picks one. The real problem is that signals get split across several URLs, and a different version than the one you want can end up in search results.

Source of duplicationExampleFix
Product in several categories/shoes/running/nike-pegasus and /sale/nike-pegasusOne product URL independent of the category, otherwise a canonical tag or 301
Sorting and filters/shoes/?order=pricenoindex or a robots.txt disallow (not both at once)
Tracking parameters/shoes/?utm_source=newsletter, ?gclid=Canonical to the clean URL, never link to them internally
Protocol, www and trailing slashhttp:// vs https://, /shoes vs /shoes/One version, 301 from all others
Letter case/Shoes/ vs /shoes/Pick one case, redirect or return 404 for the rest
Product variants/t-shirt?color=redIts own URL per variant, canonical without the optional parameter
Manufacturer descriptionsThe same text on dozens of storesYour own copy, at least for best sellers
Language versionsA German store copied to /at/ with no changeshreflang, and a canonical in the same language
The fixes follow Google’s documentation listed in the sources. The exact setup depends on your platform.

For products that sit in several categories, the cleanest option is a product URL without the category path. There is nothing to canonicalize, and you can move products between categories without redirects.

Manufacturer descriptions won’t get you penalized, but they won’t help you either. In its 2008 post, Google mentions Amazon affiliates who struggle to rank with content taken straight from Amazon and asks how they expect to outrank Amazon with the exact same listing. A store with manufacturer copy faces the same problem against every bigger competitor using that text. Write your own descriptions covering what shoppers actually care about (dimensions, differences between models, common reasons for returns), at least for the products that bring in most of your revenue.

For variants, decide whether each color or size gets its own page. Google recommends a separate URL for each variant, either as a path segment or a query parameter, and using the URL without the optional parameter as the canonical. Variants distinguished only by a fragment after # count as a single page.

06Canonical, noindex, robots.txt: which one when

Three tools, three different jobs. Problems start when they get mixed up or stacked.

What each tool does: robots.txt controls crawling, not indexing; noindex removes a page from the index but the page must be crawlable; canonical is a hint search engines may ignore; use a 301 when retiring a URL; 404 and 410 remove the page; search engines generally ignore the part after #
Each tool solves a different problem. Most mistakes come from combining them.

A canonical tag is a hint, not a rule. Google says so in exactly those words and may pick a different URL. Its troubleshooting guide notes that re-evaluation can take up to two weeks. Sitemaps are an even weaker signal: Google ranks redirects and canonical tags as strong signals and sitemap inclusion as a weak one.

robots.txt does not control indexing. A disallowed URL can still appear in search results, just without a description, if other pages link to it. Search Console then lists it as “Indexed, though blocked by robots.txt.”

noindex behind a robots.txt block does not work. The crawler never fetches the page, so it never sees the tag. Google spells this out. Yet it is a common “just to be safe” setup. Shoptet’s demo store, for instance, adds noindex to the ?order=price sort URL and disallows it in robots.txt at the same time. If no external links point to those URLs, it usually doesn’t matter. But if they are already indexed, the robots.txt block keeps them there.

noindex plus a canonical pointing elsewhere sends mixed signals. One says “don’t index me”; the other says “pass my signals to that page.” Google’s canonicalization guide explicitly advises against using noindex to choose a canonical. Pick one.

If you’re still looking for parameter settings in Search Console, they’re gone. Google retired the URL Parameters tool in 2022, noting that only about 1% of the configurations were useful for crawling. The replacements are robots.txt rules and well-designed URLs.

07Out-of-stock products, discontinued items and empty categories

A temporarily sold-out product should keep its page. Google advises leaving the page up, marking the item as out of stock and updating the structured data. Don’t use the Removals tool for this. Shoppers see what is going on, and the page keeps its rankings for when the product is back.

Product lifecycle and SEO: in stock returns 200 with InStock availability, temporarily sold out keeps the page with OutOfStock, discontinued returns 404 or 410 or a 301 to a genuine replacement, empty category gets noindex or 404, never an empty page with a 200 status
What to do with product and category pages at each stage.

A permanently discontinued product should return 404 or 410. Google treats all 4xx errors except 429 the same way, so the claim that 410 works faster in Google is not supported by its documentation. A 301 redirect makes sense only when there is a genuine replacement, such as the next model in the same line. Mass redirects to the homepage don’t help.

An empty category should get a noindex tag according to Google, or a 404 if your store removes it from navigation anyway. An empty page with a 200 status shows up in Search Console as a soft 404, and according to the crawl budget documentation Google will keep crawling it.

08Structured data for products

Google distinguishes two kinds of product markup. Product snippets are for pages where people can’t buy the product directly, such as review sites. Online stores should use merchant listing markup, which requires a name, an image and an Offer with a price greater than zero and a currency. If you provide the merchant listing properties, your pages are generally eligible for product snippets as well.

Review markup has to match what is visible on the page. In July 2026, Google also added an explicit guideline against fake and undisclosed incentivized reviews to its review snippet documentation. National consumer laws in the EU add their own rules on top. In Czechia and Slovakia, for example, a store that shows reviews must say whether and how it checks that they come from real customers.

09Language and country versions

Many stores sell in several countries. Different language versions are not duplicates as long as the main content is translated. Google puts it precisely: language versions count as duplicates only if the primary content is in the same language, meaning only the header, footer and navigation were translated. Regional versions in the same language, such as a German store for Germany and Austria, are duplicates, and Google recommends handling them with both canonical tags and hreflang.

  • hreflang in both directions: each version must list itself and all other versions, otherwise Google may ignore the annotations. It works across domains, too.
  • Correct codes: language first (ISO 639-1), then an optional region, so de-AT or en-GB. A country code on its own is not valid, which is why cz for Czech is a common error (the language code is cs).
  • Canonical in the same language: a German page should not point its canonical to the English one. Google says so explicitly.
  • Unreviewed machine translation: Google’s spam policies on scaled content list automated translation among the examples when it adds little value for users. Translation as such is not banned.

10How to audit your store in an afternoon

In Google Search Console, open the Page indexing report and go through the reasons pages are not indexed. For an online store, these statuses matter most:

  • Discovered, currently not indexed: Google knows the URL but postponed crawling. Hundreds or thousands of product URLs with this status mean the crawler is spending its time elsewhere.
  • Duplicate, Google chose different canonical than user: your canonical tag was not accepted. Check whether the pages differ too much, or whether internal links point to another version.
  • Indexed, though blocked by robots.txt: URLs you blocked too late. To get them out of the index, allow crawling temporarily and add noindex.
  • Soft 404: empty categories, nonexistent pagination, filters with no results.

Then look at one category the way a crawler would:

  • Open it with a sort order, one filter, two filters and ?utm_source=test. Does each version have the right robots tag and canonical?
  • Request a pagination page that does not exist, say page 99. Does the store return a 404?
  • Try the URL with http://, without the trailing slash and in uppercase. Does each end in a single redirect to the right version?
  • Do “Add to cart” and “Add to wishlist” links point to URLs with parameters? If so, turn them into buttons or disallow those URLs in robots.txt.
  • Does the sitemap list only URLs that return 200 and have a self-referencing canonical?

If the audit shows that your current platform will not let you control filters and URLs, read our comparison of Shopify and custom ecommerce. We plan a clean category structure from the start when we build custom online stores.

11Frequently asked questions

Does Google penalize online stores for duplicate content?

Not for ordinary duplication within a store. Google groups duplicate URLs and shows one of them. The downside is that signals get split and a different version than you want may rank. Penalties apply to copying other sites’ content to manipulate rankings.

Should I let search engines index my filter pages?

Only the combinations people search for that have enough products and their own heading and intro. Keep the other filters for shoppers and close them to search engines, ideally with a robots.txt disallow or by putting filters after # in the URL.

Is a canonical tag to the category enough for filter pages?

For a smaller store it often is, but it is only a hint. Crawlers keep fetching the filtered pages, and Google may not accept the canonical. For large catalogs, Google recommends blocking filter URLs in robots.txt.

Can I use noindex and block the page in robots.txt at the same time?

There’s no point. The crawler won’t fetch a blocked page, so it never sees the noindex tag. Choose one: noindex to remove a page from the index, robots.txt to stop it from being crawled.

What should I do with an out-of-stock product?

If it is temporarily sold out, keep the page and mark it as out of stock, including in structured data. If it is discontinued for good, return a 404 or 410, or redirect to a genuine replacement.

Is FAQ and breadcrumb markup still worth it for ecommerce SEO?

Breadcrumb markup, yes: it still shows on desktop and helps Google understand your structure. The FAQ rich result disappeared on May 7, 2026, so FAQ markup no longer changes how your result looks.

12Sources

Data as of September 30, 2026. We tested the Shoptet demo store in its default configuration on September 30, 2026. Individual stores may be set up differently.

LISTIFY teamWebsites, apps and marketing from Prague since 2008

More articles

All articles →
SEOSeptember 28, 2026 · 19 min read

SEO in 2026: What Actually Gets You to the First Page of Google

SEOSeptember 28, 2026 · 26 min read

The 3-Second Rule: Why the First Moments Make or Break Your Website and App

App developmentSeptember 30, 2026 · 21 min read

User Onboarding: How to Stop Losing New Users in the First Few Minutes

Share this page

By email

Got an idea?

On a short call, we'll find out what you need and suggest the next step. Then you'll get a proposal with a fixed price and a timeline.

+420 771 166 199Mon to Fri, 8:30 a.m. to 4:00 p.m. (Prague time) · info@listify.cool

When should we call you?

Pick a day and a time window. We'll call you, and it takes about 15 minutes.

Day