Duplicate Content: How to Find It and Fix It Before Review
The two kinds of duplicate content that trigger low-value rejections, how to detect them on your own site, and how to fix each properly.
You can have flawless grammar, a clean design, and a fast site, and still get rejected for one reason reviewers rarely spell out: your content isn’t original. Duplicate and copied content is one of the fastest ways to trip AdSense’s low-value-content filter, because it signals that your site adds nothing the web doesn’t already have. The tricky part is that “duplicate” comes in two flavors: content you took from other sites, and content that repeats itself across your own pages. Both hurt you, and most rejected applicants have some of each without realizing it. This guide shows you how to find every instance and fix it properly before you hit submit.
The two kinds of duplicate content (and why both matter)
When people hear “duplicate content” they picture outright plagiarism. That’s only half the problem. AdSense reviewers care about two distinct categories, and you need to audit for both.
- Copied from other sites. This is scraping, republishing, or “spinning” someone else’s article. It’s an outright policy violation, not just a quality issue. Taking a competitor’s post and swapping a few words does not make it yours. It makes it a rewrite of their work with your name on it.
- Internal duplication on your own site. This is subtler and far more common. Thin tag, category, and date archive pages that list the same snippets. Near-identical posts targeting slightly different keywords. Boilerplate intros, disclaimers, and “about the author” blocks that repeat verbatim across dozens of pages. None of it is stolen, but it inflates your page count with pages that say nothing new.
Reviewers see both as the same underlying failure: pages that don’t justify their own existence. If you want the broader picture of what that filter is looking for, read the low-value-content fix guide alongside this one.
How reviewers and algorithms actually detect it
You’re not fooling anyone with a thesaurus. Detection today works on meaning and structure, not exact string matching.
- Phrasing and sentence-structure similarity. Modern systems compare the shape of your sentences and the sequence of ideas, not just identical words. An article that follows another’s exact paragraph order with synonyms swapped in lights up as a near-duplicate.
- Fingerprinting against the indexed web. Google already has a copy of nearly everything published. When your page enters the index, it’s compared against existing documents. If yours is the later, thinner, less-linked version, you get treated as the copy, even if you genuinely wrote it independently.
- Human spot-checks. A reviewer reading your site can feel generic, assembled-from-other-articles writing in seconds. There are no first-hand details or opinions, none of the specifics that could only come from someone who actually did the thing.
All of this targets the substance of what you wrote. Rearranging words leaves the substance identical, which is exactly why it fails.
How to check your own site before AdSense does
Do this audit before applying, not after rejection. It takes an afternoon and saves you weeks.
- Run your posts through a plagiarism checker. Paste each article (or use a bulk tool) into a plagiarism scanner and look for matches against live URLs. Anything flagged above a low percentage needs investigation, and if the match is your own content on another domain, decide which version is canonical.
- Do exact-phrase Google searches. Take a distinctive full sentence from each article, wrap it in quotation marks, and search. If other sites return that exact string, either they copied you or you copied a source. Either way, Google sees duplication.
- Use Search Console. Check the Pages report for URLs excluded as “Duplicate,” “Alternate page with proper canonical tag,” or “Crawled — currently not indexed.” A pile of excluded thin pages is Google telling you exactly which content it considers redundant. Listen to it.
Also crawl your own site the way a bot would. Look at how many URLs exist versus how many are real articles. If a 20-post blog generates 300 crawlable URLs, your archives are the problem.
Fixing internal duplication
Most internal duplication is a settings-and-structure problem, and it’s fixable in a day. Work through these in order.
- Noindex thin archive, tag, and author pages. Tag and category pages that just list post excerpts add nothing. Set them to noindex so they never enter the index as competing thin pages. Keep them for navigation if you like, but don’t let search engines rank them.
- Consolidate near-identical posts. If you have three posts covering roughly the same topic to chase three keyword variations, merge them into one strong, comprehensive article and 301-redirect the others to it. One authoritative page beats three overlapping thin ones every time.
- Use canonical tags correctly. When the same content legitimately lives at multiple URLs (print versions, filtered listings, parameter variations), point every duplicate’s canonical tag at the single preferred URL so Google consolidates them instead of splitting or penalizing them.
- Rewrite repeated boilerplate. If your intro, disclaimer, or closing paragraph is copy-pasted across every post, that repeated block dilutes each page. Trim it, vary it, or move standing legal text to a single linked page rather than stamping it on all forty articles.
The goal is simple: every indexable URL on your site should contain something a reader can’t get from another URL on your site. Pair this cleanup with a sane sense of scale: see how many articles you actually need so you’re not padding the count with filler in the first place.
Why lightly-rewritten and spun articles still fail
This is the mistake that catches ambitious applicants. You find a great article, run it through a spinner or manually swap synonyms, and tell yourself it’s “original enough.” It isn’t, for reasons that have nothing to do with getting caught.
- The value is in the substance, not the words. A spun article contains exactly the same facts, examples, and structure as the source. You’ve added zero new information to the web. That’s the definition of low value, whether or not a detector flags the phrasing.
- Synonym-swapping degrades quality. Spun text reads slightly off, with awkward word choices and stilted transitions. Reviewers notice, and it actively signals manipulation.
- You inherit someone else’s errors and gaps. If the source was wrong or outdated, so are you, because you never actually understood the material well enough to write it yourself.
The fix isn’t a better spinner. The fix is original first-hand value: your own testing, your own screenshots and numbers, your own opinion about what works and what doesn’t. Write the thing only you could write, the post that includes the detail you learned by actually doing it. That’s the content no algorithm flags and no reviewer rejects. For the exact policy language behind all of this, keep the AdSense content policy guide handy.
Your pre-submission audit checklist
Run this list before you apply. If you can’t check every box, you’re not ready.
- Every article passes a plagiarism check against live URLs, with no unexplained external matches.
- Exact-phrase searches of your distinctive sentences return only your own site.
- Search Console shows no meaningful pile of “Duplicate” or “Alternate page” exclusions.
- Thin tag, category, author, and date archives are set to noindex.
- Near-identical posts have been merged and redirected to a single canonical version.
- Canonical tags are correct on any legitimately duplicated URLs.
- Repeated boilerplate has been trimmed or moved to a single linked page.
- No article on your site is a rewrite, spin, or synonym-swap of another source. Each one contains first-hand detail only you could provide.
What it comes down to
Duplicate content fails AdSense for one honest reason: it doesn’t add anything to the web. Copying others is a policy violation, spinning is copying with extra steps, and internal duplication buries your real work under thin repetitive pages. The fix for all three is the same: original, first-hand value on every indexable URL. Clean up your archives, consolidate your overlap, throw out anything you rewrote from a source, and make sure each page earns its place. Run the checklist above, and when your content genuinely stands on its own, use the site checker tool to scan for anything you missed before you submit. Do the work once, properly, and you won’t be back here after a rejection wondering what “low value” meant.
Run the audit on your own site
See, in fifteen seconds, which of the issues described here apply to your site — and which you’ve already cleared.
Run free auditMore guides
View all →Reapplying for AdSense After Rejection: A Recovery Plan
Rejected for low value content? A calm, concrete cycle: diagnose the reason, fix the substance, wait for the re-crawl, then resubmit once.
Images for AdSense: Copyright-Safe Sources (and Why Text-Only Posts Fail)
Why text-only posts read as thin, the copyright trap of Google Images, and where to get relevant, safe visuals with proper alt text.
How to Get Real Traffic Before You Apply for AdSense
AdSense has no traffic minimum, but real human visitors help. White-hat channels — search, Pinterest and social — without ever buying traffic.