Near duplicate
Definition
A near duplicate is a page that is not an exact copy but gives substantially the same main answer as another page. Wording, sentence order, entity names, or small facts may differ, yet the useful content remains similar enough that the two URLs may not justify separate existence.
Why near duplicate matters in programmatic SEO
Near duplicates are easy to miss in programmatic SEO because structural fields can all look unique. A family can have different URLs, titles, H1s, and target keywords while reusing almost the same body, examples, proof, and recommendation on every page.
The risk is best judged across the page set. One location page may look acceptable in isolation; the repetition becomes obvious only when it is compared with ten siblings and the only consistent difference is the city name. That is why duplicate auditing is a relationship test between pages, not a checklist on one URL.
A simple example
Suppose three accounting-service pages use different service names in the title and introduction, but the pricing claim, proof points, workflow, FAQs, and recommendation are otherwise the same. They are not exact copies, yet the visitor receives almost the same answer from each page.
Shared structure alone is not the problem. An integration template can reuse the same sections safely when supported actions, authentication, setup, limitations, and examples genuinely change by integration.
Common misconception
“Reusing a template automatically makes pages near duplicates. Programmatic systems are expected to reuse structure; the concern is repeated main answers and evidence, not repeated layout.”
How pSEO Guard handles near duplicate
When body text is supplied, pSEO Guard compares planned pages for near-duplicate content and also checks patterns where pages remain almost the same after page-specific names are discounted. A verified near-duplicate finding is Blocked because the page set needs a decision before publication.
Comparison coverage remains visible. If body content is missing or a bounded comparison cannot evaluate every relevant pair, the affected rows are not treated as clean. The product threshold is an audit heuristic, not a Google requirement for a fixed percentage of unique text.