Skip to main content
Premier SEO Services

Technical SEO

Duplicate content: what Google actually does

By the Premier SEO Services teamPublished 5 min read

The penalty that does not exist

Google's documentation on duplicate content and its guidance on canonicalization both describe the same mechanism, and it is not punitive. When Google finds several URLs with the same or very similar content, it:

  1. groups them together;
  2. selects one URL as canonical, using redirects, canonical tags, internal links, sitemaps and URL tidiness as signals;
  3. shows the canonical in results and consolidates the group's signals onto it.

No ranking demotion is applied for the existence of the duplicates. Google's own documentation says duplicate content on a site is not grounds for action against that site unless it appears deceptive — scraped content republished at scale, or doorway pages spun to target many locations with one text.

What duplication does cost you is real, though, just not a penalty:

  • Split signals. Links arriving at three versions of a page are consolidated only once Google is confident they are duplicates. Until then, each URL accumulates separately.
  • Wasted crawling. Every variant is fetched before being discarded, which matters on large sites.
  • The wrong URL chosen. Google may canonicalise to the version with the tracking parameter, the http variant, or the printer-friendly page.
  • Unstable reporting. Search Console and analytics split one page's performance across several rows.

So the goal is not avoiding a penalty. It is making sure one URL per piece of content collects everything.

Where duplication actually comes from

In practice, almost all duplication on a site is generated by the site itself. The usual sources:

URL variants of one page. This is the big one, and it is invisible until you look:

https://example.com/shoes/          https://www.example.com/shoes/
http://example.com/shoes/           https://example.com/shoes
https://example.com/Shoes/          https://example.com/shoes/index.html
https://example.com/shoes/?utm_source=newsletter

That is eight URLs and one page. Each is a separate address as far as HTTP is concerned.

Parameters. Tracking parameters, session identifiers, sort orders, and filters that do not change the result set.

Pagination and archives. A blog post reachable at its own URL, plus a category page, a tag page, a date archive and an author archive, each showing the full text rather than an excerpt.

Templated pages with thin variation. Location pages where only the city name changes, or product variants with one differing attribute. These are genuinely hard cases: Google may treat them as duplicates and show one, which is reasonable when the pages really do say the same thing.

Syndication. Your article republished by a partner. Without a canonical or noindex on their copy, their version can outrank yours.

Staging and alternate hosts. A development server left crawlable, or a CDN host serving the same pages.

Localisation without hreflang. en-GB and en-US pages with near-identical text and no annotations connecting them.

To see how much of this applies to you, pick one page and request it in every variant form you can construct, checking what each returns. The redirect checker shows whether each lands on one canonical URL in a single hop, which is the outcome you want.

Fixing it in the right order

The tools are not equivalent, and the strongest one is the cheapest to apply.

1. Redirect, where one URL should simply not exist. A 301 removes the duplicate rather than annotating it. This is the right fix for protocol, host, trailing slash and letter-case variants. Enforce it in one hop: http to https, non-preferred host to preferred, and the trailing-slash convention, all resolved by a single rule rather than three chained ones.

2. Self-referencing canonical tags everywhere. Every page states its own preferred URL in absolute form. This handles parameters you cannot control, such as a tracking parameter someone appends to a shared link. The canonical URL generator normalises an address into the exact form to use in the tag, in internal links and in the sitemap.

3. Consistent internal linking. Link to the canonical form everywhere, including in navigation, breadcrumbs, sitemaps and structured data. Internal links are a canonicalization signal, and a site that links to three variants is voting against itself.

4. Block or `noindex` where the variant has no value to anyone. Disallow sort parameters to save crawling; noindex internal search results and thin tag archives.

5. Consolidate, where two pages genuinely compete. If two articles target the same question, the honest fix is to merge them into the better one and redirect the weaker URL. Two half-pages rarely beat one complete page, and the keyword density checker run over both quickly shows how far their topical overlap really goes.

6. For syndicated copies, ask partners for a canonical pointing to your original, or a noindex. Publishing the original first and getting it indexed first also helps.

What not to reach for: rewriting text purely to make it "unique" when the duplication is a URL problem. Spinning near-identical variants of a location page is the behaviour Google's doorway-page guidance describes, and it addresses the wrong cause.

Checking your work

A short audit that catches most duplication:

Confirm one canonical host and protocol. Request all four combinations of http/https and with/without www and check each ends at the same URL after one redirect.

Check the trailing-slash convention. Both /page and /page/ should exist as one canonical form, with the other redirecting. Serving both with 200 is a duplicate pair on every page of the site.

Add a parameter and see what happens. Request your homepage with ?x=1 appended. If it returns 200 with a canonical tag pointing at the clean URL, you are fine. If the canonical includes the parameter, every shared link with tracking creates a new indexable URL.

Look for index filenames. / and /index.html or /index.php both returning 200 is a duplicate pair.

Read the tags a page really sends with the meta tags analyzer, because a plugin can override the canonical your template sets.

Use Search Console. The Page indexing report distinguishes the outcomes: "Alternate page with proper canonical tag" means the mechanism is working as intended; "Duplicate without user-selected canonical" means you have not stated a preference; "Duplicate, Google chose different canonical than user" means your other signals contradict your tag.

A note on scale, to keep this in proportion. A handful of parameter variants on a small site is not worth a project. The cases that genuinely repay the work are sites where one page exists at many addresses systematically, because the waste multiplies by the number of pages: a trailing-slash inconsistency on a 50,000-page site is 50,000 duplicate pairs.

Sources