Teaching an Agent to Read a Form It Has Never Seen
Field mapping starts as a similarity search against forms we already solved, and only falls back to inference. The embedding step pays for itself by the second campaign.
There is no standard for a directory submission form. One asks for a site name, a URL and a category. The next wants a company description in two lengths, a contact title, a logo at an exact aspect ratio, and a checkbox whose label is a sentence long. The field names in the markup are frequently worse than the labels, and sometimes both are in a language the campaign is not.
The obvious approach is to hand the whole form to a language model on every attempt and ask it what goes where. It works. It is also slow, metered per publisher, and non-deterministic in a place where determinism is worth a lot — the same form should not map differently on Tuesday than it did on Monday.
Similarity first, inference second
Before any model is asked anything, the agent builds a fingerprint of the form: the visible labels, the input names and placeholders, the control types, the required flags, and the order they appear in. That fingerprint is embedded and searched against every form the tenant has already solved. Directory software is not infinitely varied — a large share of the long tail is the same handful of listing platforms with a different logo on top.
The cheapest field mapping is the one you already solved. Inference is the fallback path, not the default one.
Three bands, three behaviours
The nearest-neighbour distance decides what happens next. The bands are deliberately coarse, because a mapping that is nearly right is more dangerous than one that is obviously wrong.
Reuse applies the stored mapping directly and asserts that every field it expects is still present. Adapt applies the stored mapping to the fields it recognises and sends only the leftovers to the model, which is usually one or two additions after a platform update. Infer builds the mapping from scratch and, if the submission succeeds and verifies, writes the result back so the next publisher on that platform lands in the reuse band.
Inference is constrained, not open-ended
When the model is called, it is not asked to fill in a form. It is asked to produce a mapping between the fields it can see and a closed vocabulary of campaign values: site name, target URL, contact email, category, short description, long description, logo, keywords, and a small number of others. It returns field identifiers, never prose.
Anything outside that vocabulary is left empty and the submission is flagged for review. This is the single most important rule in the stage. A model given a required field it does not understand will invent a plausible value, and a plausible value is how you end up with a live listing whose company description is a paraphrase of the form's own help text.
Why it pays for itself by the second campaign
The first campaign in a vertical is mostly inference and is priced accordingly. The second one in the same vertical draws on the mappings the first one earned, and the inference calls drop to the genuinely novel forms. The embedding index is the asset — it accumulates, it is per tenant, and it makes each campaign cheaper than the one before it rather than the same price forever.
That is also why the write-back is gated on verification rather than on submission. A mapping is only worth caching if the listing it produced actually appeared. Caching a mapping that submitted cleanly and published nothing propagates the failure across every future campaign that matches it.
What still breaks
Multi-step wizards, where fields on step three depend on the category chosen on step one. Forms that render additional required inputs only after a select changes. Honeypot fields that are invisible to a person and perfectly visible to a headless browser, which the agent must learn to leave alone. And forms whose required markers are decorative, so the real validation only announces itself on submit.
None of these are solved by a better prompt. They are solved by treating form filling as an interaction with a state machine rather than as a single mapping problem, which is the direction the stage is heading.