Unique at Scale: Semantic De-duplication of Generated Copy
Five hundred descriptions in one category will collapse toward each other unless you measure it. The similarity threshold we settled on, and how we enforce it pre-write.
Ask a language model for a company description five hundred times, for the same company, across the same category of directories, and you will not get five hundred descriptions. You will get about twelve, each with several dozen paraphrases wearing different verbs.
The problem is not plagiarism and no exact-match check will catch it. It is convergence: the model has a strong prior for what this kind of copy sounds like, every generation drifts toward that prior, and the results are individually fine and collectively a footprint. A directory moderator who has seen four of your listings can recognise the fifth.
Measure before you write, not after
The de-duplication check runs between generation and persistence, never as a clean-up pass. The candidate description is embedded, searched against the tenant's existing copy in the same vertical, and compared against the nearest neighbour's cosine similarity. Only then does it get written and queued for submission.
Running it afterwards is the tempting design, because it is easier to build as a report. It is also useless: by the time you know two descriptions are near-identical, both are live on publishers you cannot edit.
Reject discards the candidate and regenerates from scratch with the neighbour supplied as an explicit constraint. Rewrite keeps the candidate's structure and asks for a targeted change to the parts that overlap. Accept writes it and adds it to the index, which means the next candidate has one more thing to be different from — the constraint tightens on its own as the campaign runs.
Where the threshold came from
Not from a paper. Set the cut-off too strict and the regeneration loop turns into thesaurus soup, where the model has exhausted its good phrasings and is now producing awkward ones purely to satisfy a distance metric. Set it too loose and you ship templates with the nouns swapped.
Similarity thresholds cannot be chosen from a paper. Print fifty pairs at each candidate cut-off and read them.
That is the entire methodology and it is worth the afternoon. The numbers above are where our own reading landed for short marketing copy in English; they are not a universal constant, and they will move for longer bodies or other languages. What generalises is the procedure, not the constant.
Regeneration needs a different prompt, not a different seed
Re-rolling the same prompt and hoping for variance is the most common mistake here, and it mostly produces another sample from the same cluster you were trying to escape. The rewrite call receives the nearest neighbour verbatim as an avoid-set, plus a required angle the new version has to take — a different opening, a different attribute of the product foregrounded, a different sentence structure.
The loop is bounded at three attempts. If the third candidate is still inside the band, the item goes to the operator with all three versions shown, because at that point the constraint is more likely to be the campaign's input material than the model's behaviour.
Scope the comparison correctly
The comparison set is per tenant and per vertical, and both halves of that matter. Comparing across tenants is a data-isolation problem before it is anything else, and it is also wrong on the merits — two unrelated clients are allowed to have similar descriptions. Comparing against a global corpus is the opposite failure: everything is somewhat similar to everything, the distances compress, and the threshold stops discriminating.
The vectors live in per-tenant collections for exactly this reason. It also means the index is a per-customer asset that gets more useful the longer they run campaigns, rather than a shared pool that gets noisier.
What we still get wrong
Uniqueness is not quality. A description can clear every threshold in the table and still be a bad piece of copy, and the distance metric has no opinion about that. We measure convergence because it is measurable; the judgement about whether the writing is any good remains a sampled human review, and we have not found a way around that which we trust.