Glossary · Quality & Auditing

Existing-site collision

Definition

An existing-site collision happens when a proposed page conflicts with something already live on the site. The overlap can be an exact URL, a repeated title, substantially similar body content, or known target intent when the comparison source actually provides that intent data.

Why existing-site collision matters in programmatic SEO

A new page family can be perfectly clean against itself and still recreate an older service page, guide, category, landing page, or previous programmatic batch. The generator only knows the rows it was given; the domain already has a history that the new matrix may not reflect.

Checking the existing site therefore answers a different question from checking the planned set. Internal uniqueness asks whether proposed rows collide with each other. Existing-site review asks whether the settled plan is trying to create a second owner for something that already exists.

A simple example

A new matrix proposes `/services/accounting/`. The URL is unique inside the batch, but the planned body closely repeats an established `/small-business-accounting/` page that already serves the same visitor need. The new row is structurally clean and still creates a live-site collision.

A different row may use a new URL and body but repeat the exact title of an existing page. That is weaker evidence than an exact URL or body duplicate, but it still deserves review because it can reveal unclear page ownership.

Common misconception

“If every proposed row is unique against every other proposed row, the page set is safe. The existing site is a separate source of duplication, ownership, and canonicalization risk.”

How pSEO Guard handles existing-site collision

pSEO Guard's bounded Existing Site Guard compares the planned set with pages its crawl actually retrieved. An exact live URL collision is Blocked, a live title collision is Review, and a planned body that is a near duplicate of retrieved live content is Blocked.

Coverage is part of the result. Pages the crawl did not retrieve cannot be treated as checked, and a normal crawl cannot reconstruct a live page's historical target query. Intent is therefore reported as not measured unless the source provides it instead of being guessed from the page.

Go deeper

Related articlePrevent Keyword Cannibalization Before PublishingDocumentationExisting Site ScanTry itpSEO Audit