# Duplicate URLs and Which One Google Should Keep

- Source: https://simeoncreatives.com/blog/duplicate-urls-which-one-google-should-keep
- Hub: SEO, GEO & Digital Marketing
- Author: Simeon Matheka, Founder & Creative Director
- Published: 2026-09-28
- Updated: 2026-09-28
- Reading time: 18 min

HTTP and HTTPS, www and the bare host, a trailing slash, a campaign parameter. Pick the URL Google should keep, and line up redirects, rel=canonical, and the sitemap so they do not argue.

The same page answers on more than one address. HTTP and HTTPS. www and the bare host. A trailing slash and no trailing slash. A clean path and the same path with a campaign parameter. You want one URL in Search. Google can choose without your help. You should still tell it, because the signals only work if they agree.

This is the ongoing duplicate problem. If every URL on the site just changed, that indexation job is [search migration: canonicals, sitemaps, and indexation](https://simeoncreatives.com/blog/search-migration-canonicals-sitemaps-indexation). Do not run a domain move from this page.

## You can skip a canonical, until the duplicates start costing you

Sourced from Google. If you do not specify a canonical URL, Google will identify which version is objectively the best one to show. None of the methods below are required. There are still reasons to state a preference. People should land on the URL you want in results. Signals such as links can consolidate onto that URL. Metrics are easier when one piece of content is not split across addresses. And you would rather Googlebot spend time on new or updated pages than on duplicate copies. Source: [How to specify a canonical URL with rel="canonical" and other methods](https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls) (Search Central, page last updated 2026-07-10 UTC, checked 28 September 2026).

The dress-page example on that document is the shape of a parameter duplicate. You would rather people reach https://www.example.com/dresses/green/green-dress.html than https://example.com/dresses/cocktail?gclid=ABCD. Links that point at the parameter URL can consolidate onto the preferred URL if that preferred URL becomes canonical. Tracking one product, or one article, gets harder when the reports are split across those strings. That is the job. Not a new domain.

## Signal strength, in the order Google wrote it

The document lists methods in order of how strongly they can influence canonicalization. It names two strong signals and one weak signal. Methods can stack. Using two or more increases the chance that your preferred URL is the one that appears in results. Stacking only helps when they name the same URL.

| Method | Strength on that page | Use it when |
| --- | --- | --- |
| Redirects | Strong signal that the target should become canonical | You are deprecating a duplicate. All permanent redirect methods have the same effect on Google Search. How fast they are noticed can differ. Server-side HTTP redirects are the quickest. |
| rel=canonical link annotations | Strong signal that the specified URL should become canonical | The HTML link element, or the HTTP header. Pick one. Using both is supported and easier to get wrong. |
| Sitemap inclusion | Weak signal that listed URLs should become canonical | A large site where you want a simple list of preferred URLs. Google still has to decide which other URLs are duplicates, based on similarity. |

The comparison table on the same page matches that ranking. A sitemap is simpler to maintain on a large site, and it is a less powerful signal than the rel=canonical mapping. Google still has to figure out which URLs are the duplicates of a URL you listed. Redirects are for getting rid of an existing duplicate, not for a URL you still want people to open.

> Framework: a sitemap does not outvote a redirect or a rel=canonical that points somewhere else. Those two are the strong signals. The sitemap is the weak one. Make the strong ones name the same URL, and put that URL in the sitemap too.

## Conflicting signals throw the preference away

Do not specify different URLs as canonical for the same page with different techniques. The example on the page is one URL in the sitemap and a different URL in rel=canonical. Do include a rel=canonical link on the canonical page itself, the self-referential canonical. Do not use a URL fragment as the canonical. Google generally does not support URL fragments for this.

- **robots.txt: **Do not use it for canonicalization. Google may still index a disallowed URL without its content. A block is not a preference.
- **URL removal: **Do not use the removal tool to pick a winner. It hides all versions of a URL from Search.
- **noindex: **Not the tool for choosing a canonical inside one site. It blocks the page from Search. The link annotation is the preferred fix.
- **Internal links: **Link to the canonical URL, not to a duplicate. Consistent internal links help Google see the preference.
- **hreflang: **If you use hreflang, the canonical should be in the same language, or the best substitute if a same-language canonical does not exist. For localization, Google prefers URLs that sit in an hreflang cluster over a URL that the cluster does not mention.

rel=canonical annotations that also carry hreflang, lang, media, or type are ignored for canonicalization. Alternate versions get their own link annotations, such as rel="alternate" with hreflang. Google supports explicit canonical annotations as described in RFC 6596. That is the spec the help page names. It is not a reason to invent extra attributes on the canonical link.

## HTTP, www, parameters, and the trailing slash

Pick one preferred URL and point every strong signal at it. For protocol, Google prefers HTTPS pages over equivalent HTTP pages as canonical, unless something conflicts. Add a redirect from HTTP to HTTPS, or a rel=canonical link from the HTTP page to the HTTPS page, or implement HSTS. Avoid a bad certificate and avoid sending HTTPS users to or through HTTP. Those make Google prefer HTTP very strongly, and HSTS cannot override that. Do not list the HTTP URLs in the sitemap or in hreflang instead of the HTTPS URLs. The certificate has to match the site URL, or be a wildcard that covers the subdomains you actually serve. A certificate for subdomain.example.com served on example.com is the kind of mismatch the page tells you to avoid.

For hostnames, the redirect section uses three ways to reach one home: https://example.com/home, https://home.example.com, and https://www.example.com. Choose one. Redirect the others to it with a server-side permanent redirect if you are deprecating those duplicates. The www choice is that kind of duplicate. It is not, by itself, a domain migration.

For parameters, the green-dress URL versus the cocktail URL with gclid is the worked example. The clean URL is the one you want people to see. The parameter URL can still be crawled. Your canonical, your redirects if you are willing to drop that entry URL, and your sitemap should agree on the clean one. Do not put only the parameter URL in the sitemap while the link element points at the clean URL.

A trailing slash is the same class of problem: two strings, one page. The canonical document’s examples are protocol, host, and parameters. It does not choose slash or no slash. Choose one form. Redirect the other if you are deprecating it. Put the chosen form in the self-referential canonical and in the sitemap. Internal links should use that form. If the slash URLs and the no-slash URLs both return 200 with the same content and point their canonicals at themselves, you have two preferences, which is no preference.

## HTML uses the link element. PDF uses the header.

A rel=canonical link element goes in the head of the HTML. It has to appear in the head, and at least the head section needs to be valid HTML. Use an absolute URL. Relative paths are supported and not recommended. They cause long-term mistakes, including a testing host that gets crawled by accident. Add the same self-referential element on the canonical page. If you also have a separate mobile URL, the canonical page can point at it with rel="alternate", and still canonical itself.

#### HTML link element

```html
<link rel="canonical" href="https://example.com/dresses/green-dresses" />
```

#### HTTP header

```http
Link: <https://www.example.com/downloads/white-paper.pdf>; rel="canonical"
```

*Absolute URLs. One method, not both with different targets. The header is how a non-HTML file states a canonical.*

The link element only works for HTML pages, not for files such as PDF. The HTTP header is the alternative the document names, including for non-HTML documents. If the PDF and a Word file each have their own URL, you can send the header on the file that should not be canonical so Googlebot knows which URL is. The example on the page puts the header on the .docx version and points at the PDF. Google supports this header method for web search results only. Choose the element or the header and stay with it. Both at once is supported, and it is how you accidentally publish one URL in the header and another in the element.

If the page is built with JavaScript, the canonical has to be obvious. The best way is to put the canonical URL in the HTML source and make sure JavaScript does not change that link element. If you cannot set it in the HTML source, leave it out of the HTML and set it only with JavaScript. Do not set one in the source and a different one after render. If the first HTML is empty, start with [why crawlers saw an empty React site](https://simeoncreatives.com/blog/why-crawlers-saw-an-empty-react-site), then confirm the canonical with [URL Inspection](https://simeoncreatives.com/blog/search-console-url-inspection-javascript-site).

## A content management system can write the tag. You still pick the URL

If you use a content management system (CMS) such as WordPress, Wix, or Blogger, you may not edit the head by hand. The document says the CMS may have a search setting, or another control, that writes the canonical. Search that CMS’s instructions for changing the head. The decision is still yours: one preferred URL, strong signals aimed at it, sitemap included, no second canonical hiding in a plugin.

On a large site the link-element map is harder to maintain, especially when URLs change often. That is a maintenance con on Google’s comparison table, not a reason to use only the weak signal. A sitemap scales more easily and still loses when it disagrees with a redirect or a rel=canonical. Use the sitemap as the list of what you think is important. Use redirects when a duplicate should stop existing as its own URL. Use the link element or the header to name the keeper on the pages that remain.

## Filled hypothetical: four addresses, one article

Hypothetical, labeled as such. One article loads on http and https, on www and the bare host, with and without a trailing slash, and with a gclid parameter from an ad. The sitemap lists the HTTP www URL because that is what an old plugin exported. The HTTPS page’s canonical points at the HTTPS non-www URL without a slash. Internal navigation sometimes adds the slash. Nothing redirects.

Google may still pick a URL. You do not control which one, and the signals you did send disagree. The fix is not another sitemap row. Choose the HTTPS host and the slash form. Redirect the HTTP URLs and the other host to it. Put a self-referential canonical on that URL. Point every duplicate HTML page’s canonical at it. List that URL, not the HTTP one, in the sitemap. Link internally only to it. Then inspect one duplicate and the keeper. If they still disagree after render, the JavaScript is changing the element, and the source-order rule above applies.

## What this does not prove

A strong signal is not a guarantee. The page says these methods influence canonicalization and that combining them raises the chance your URL is shown. It does not say Google must obey. Sitemap inclusion does not become strong because the file is large. noindex does not “keep the duplicate out of the cluster” while leaving it useful. And a trailing-slash preference is your rule, lined up with the signals, not a ranking factor this document states.

If the templates cannot hold one canonical in the head, that is a [website](https://simeoncreatives.com/websites) problem. [Start a conversation](https://simeoncreatives.com/contact) with the four URLs that currently 200, and with which one you want kept.

## FAQs

### Is a sitemap enough to choose the canonical URL?

No. Google calls sitemap inclusion a weak signal that the listed URLs should become canonical. Redirects are a strong signal that the target should become canonical. A rel=canonical link annotation is also a strong signal. If the sitemap names one URL and rel=canonical names another, those signals conflict. Google says not to do that.

### Does the rel=canonical link element work on a PDF?

The link element works for HTML pages, not for files such as PDF. For a PDF or another non-HTML file, Google names the rel=canonical HTTP header as the alternative. Google supports that header method for web search results only. Use an absolute URL in the header, same as in the link element.

### Should I noindex the duplicate instead?

Google does not recommend noindex to stop a duplicate from being chosen as canonical inside one site, because noindex blocks the page from Search entirely. rel=canonical link annotations are the preferred approach. Do not use robots.txt to canonicalize either. Google may still index a URL disallowed in robots.txt, without its content. Do not use the URL removal tool. It hides all versions of a URL from Search.

### Why might Google keep the HTTP URL instead of HTTPS?

Google prefers HTTPS over equivalent HTTP pages, except when there are issues or conflicting signals. The page lists an invalid certificate, insecure dependencies other than images, an HTTPS page that redirects users to or through an HTTP page, and a rel=canonical link from the HTTPS page to the HTTP page. Bad certificates and HTTPS-to-HTTP redirects make Google prefer HTTP very strongly. HSTS does not override that strong preference.

### Is this the guide for a domain migration?

No. This page is for duplicates that stay on the site: protocol, host, trailing slash, parameters. A one-time move where every URL changes, including the indexation work in Search Console, is the search migration guide. If the URLs are not changing and only the host or DNS is, that is a third document on Search Central, and the migration article points at it.
