"Duplicate content penalty" is mostly a myth, but duplicate URLs are a real problem: they split signals, waste crawl budget, and confuse which page should rank.
Where it comes from
- The same page reachable at http and https, or with and without www
- Tracking parameters (?utm_source=…) creating many URLs
- Pagination, filters and sorting on e-commerce sites
- Print versions and session IDs
- Syndicated articles published on multiple sites
What it costs
The wrong URL can rank, links can point at three copies instead of one, and crawlers can waste time on near-identical pages instead of your new content.
The fixes
1. Canonical tags — point every variant to the preferred URL. See canonical tags explained.
2. Consistent internal links — always link to the canonical version.
3. Redirects — 301 alternate hosts (e.g. non-www to www).
4. robots.txt — block infinite URL spaces like internal search results.
5. noindex — for thin variants you cannot canonicalise cleanly.
Check a page
Paste a page's HTML into the free canonical URL checker to confirm its canonical, title and description are what you expect.