Agent-friendly documentation: a Markdown index for AI agents is at /llms.txt. Every page is also available as clean Markdown by appending .md to its URL (e.g. /pricing.md), or by requesting it with an Accept: text/markdown header.

BuildSEOAI
All posts
Engineering· 7 min read

Unique at Scale: Semantic De-duplication of Generated Copy

Five hundred descriptions in one category will collapse toward each other unless you measure it. The similarity threshold we settled on, and how we enforce it pre-write.

AP
Anurag Pattnaik
Platform engineering, Andolasoft · June 3, 2026

Ask a language model for a company description five hundred times, for the same company, across the same category of directories, and you will not get five hundred descriptions. You will get about twelve, each with several dozen paraphrases wearing different verbs.

The problem is not plagiarism and no exact-match check will catch it. It is convergence: the model has a strong prior for what this kind of copy sounds like, every generation drifts toward that prior, and the results are individually fine and collectively a footprint. A directory moderator who has seen four of your listings can recognise the fifth.

Measure before you write, not after

The de-duplication check runs between generation and persistence, never as a clean-up pass. The candidate description is embedded, searched against the tenant's existing copy in the same vertical, and compared against the nearest neighbour's cosine similarity. Only then does it get written and queued for submission.

Running it afterwards is the tempting design, because it is easier to build as a report. It is also useless: by the time you know two descriptions are near-identical, both are live on publishers you cannot edit.

SIMILARITY
ACTION
0.94 and above
Reject
0.86 to 0.93
Rewrite
Below 0.86
Accept
No neighbours
Accept

Reject discards the candidate and regenerates from scratch with the neighbour supplied as an explicit constraint. Rewrite keeps the candidate's structure and asks for a targeted change to the parts that overlap. Accept writes it and adds it to the index, which means the next candidate has one more thing to be different from — the constraint tightens on its own as the campaign runs.

Where the threshold came from

Not from a paper. Set the cut-off too strict and the regeneration loop turns into thesaurus soup, where the model has exhausted its good phrasings and is now producing awkward ones purely to satisfy a distance metric. Set it too loose and you ship templates with the nouns swapped.

Similarity thresholds cannot be chosen from a paper. Print fifty pairs at each candidate cut-off and read them.

That is the entire methodology and it is worth the afternoon. The numbers above are where our own reading landed for short marketing copy in English; they are not a universal constant, and they will move for longer bodies or other languages. What generalises is the procedure, not the constant.

Regeneration needs a different prompt, not a different seed

Re-rolling the same prompt and hoping for variance is the most common mistake here, and it mostly produces another sample from the same cluster you were trying to escape. The rewrite call receives the nearest neighbour verbatim as an avoid-set, plus a required angle the new version has to take — a different opening, a different attribute of the product foregrounded, a different sentence structure.

The loop is bounded at three attempts. If the third candidate is still inside the band, the item goes to the operator with all three versions shown, because at that point the constraint is more likely to be the campaign's input material than the model's behaviour.

Scope the comparison correctly

The comparison set is per tenant and per vertical, and both halves of that matter. Comparing across tenants is a data-isolation problem before it is anything else, and it is also wrong on the merits — two unrelated clients are allowed to have similar descriptions. Comparing against a global corpus is the opposite failure: everything is somewhat similar to everything, the distances compress, and the threshold stops discriminating.

The vectors live in per-tenant collections for exactly this reason. It also means the index is a per-customer asset that gets more useful the longer they run campaigns, rather than a shared pool that gets noisier.

What we still get wrong

Uniqueness is not quality. A description can clear every threshold in the table and still be a bad piece of copy, and the distance metric has no opinion about that. We measure convergence because it is measurable; the judgement about whether the writing is any good remains a sampled human review, and we have not found a way around that which we trust.

TAKEAWAYS
Check similarity between generation and persistence — a de-duplication report on published copy is a post-mortem, not a control.
Use three bands rather than one threshold so near-misses get a targeted rewrite instead of a full re-roll.
Pick the cut-off by reading pairs at each candidate value; the procedure generalises, the constant does not.
Feed the nearest neighbour back into the regeneration prompt as an avoid-set — re-rolling the same prompt resamples the same cluster.
Scope the index per tenant and per vertical; global comparison compresses the distances until the threshold stops discriminating.
Share this post

See the Score Run on Your Category

We will qualify live publisher inventory in your vertical and show you every signal behind it.

Schedule a Demo