# Teaching an Agent to Read a Form It Has Never Seen

> Agent-friendly Markdown of https://www.buildseo.org/blog/teaching-an-agent-to-read-a-form — full index for AI agents: https://www.buildseo.org/llms.txt

_Automation · 8 min read — Anurag Pattnaik, Platform engineering, Andolasoft · July 7, 2026_

Field mapping starts as a similarity search against forms we already solved, and only falls back to inference. The embedding step pays for itself by the second campaign.

There is no standard for a directory submission form. One asks for a site name, a URL and a category. The next wants a company description in two lengths, a contact title, a logo at an exact aspect ratio, and a checkbox whose label is a sentence long. The field names in the markup are frequently worse than the labels, and sometimes both are in a language the campaign is not.

The obvious approach is to hand the whole form to a language model on every attempt and ask it what goes where. It works. It is also slow, metered per publisher, and non-deterministic in a place where determinism is worth a lot â€” the same form should not map differently on Tuesday than it did on Monday.

## Similarity first, inference second

Before any model is asked anything, the agent builds a fingerprint of the form: the visible labels, the input names and placeholders, the control types, the required flags, and the order they appear in. That fingerprint is embedded and searched against every form the tenant has already solved. Directory software is not infinitely varied â€” a large share of the long tail is the same handful of listing platforms with a different logo on top.

> The cheapest field mapping is the one you already solved. Inference is the fallback path, not the default one.

## Three bands, three behaviours

The nearest-neighbour distance decides what happens next. The bands are deliberately coarse, because a mapping that is nearly right is more dangerous than one that is obviously wrong.

| NEAREST MATCH | ACTION |
| --- | --- |
| 0.92 and above | Reuse |
| 0.75 to 0.91 | Adapt |
| Below 0.75 | Infer |
| No neighbours | Infer |

Reuse applies the stored mapping directly and asserts that every field it expects is still present. Adapt applies the stored mapping to the fields it recognises and sends only the leftovers to the model, which is usually one or two additions after a platform update. Infer builds the mapping from scratch and, if the submission succeeds and verifies, writes the result back so the next publisher on that platform lands in the reuse band.

## Inference is constrained, not open-ended

When the model is called, it is not asked to fill in a form. It is asked to produce a mapping between the fields it can see and a closed vocabulary of campaign values: site name, target URL, contact email, category, short description, long description, logo, keywords, and a small number of others. It returns field identifiers, never prose.

Anything outside that vocabulary is left empty and the submission is flagged for review. This is the single most important rule in the stage. A model given a required field it does not understand will invent a plausible value, and a plausible value is how you end up with a live listing whose company description is a paraphrase of the form's own help text.

## Why it pays for itself by the second campaign

The first campaign in a vertical is mostly inference and is priced accordingly. The second one in the same vertical draws on the mappings the first one earned, and the inference calls drop to the genuinely novel forms. The embedding index is the asset â€” it accumulates, it is per tenant, and it makes each campaign cheaper than the one before it rather than the same price forever.

That is also why the write-back is gated on verification rather than on submission. A mapping is only worth caching if the listing it produced actually appeared. Caching a mapping that submitted cleanly and published nothing propagates the failure across every future campaign that matches it.

## What still breaks

Multi-step wizards, where fields on step three depend on the category chosen on step one. Forms that render additional required inputs only after a select changes. Honeypot fields that are invisible to a person and perfectly visible to a headless browser, which the agent must learn to leave alone. And forms whose required markers are decorative, so the real validation only announces itself on submit.

None of these are solved by a better prompt. They are solved by treating form filling as an interaction with a state machine rather than as a single mapping problem, which is the direction the stage is heading.

**Takeaways**

- Fingerprint and search before you infer â€” most forms in a vertical are the same platform wearing a different theme.
- Split the match into reuse, adapt and infer bands; a nearly-right mapping is worse than an obviously wrong one.
- Constrain inference to a closed vocabulary and leave unknown fields empty rather than letting the model invent values.
- Cache a mapping only after the listing verifies, or you will propagate a silent failure across every future campaign.

## Keep Reading

- [The Hidden Cost of Manual Link Submission: A 5-Year ROI Case Study](https://www.buildseo.org/blog/the-hidden-cost-of-manual-link-submission.md): One agency spent $840k on manual link building over five years. We mapped where every dollar went, and why their cost per verified link was 12Ã— higher than they reported.
- [Cost Per Link: How To Calculate It, Why Vendors Hide It, and What It Actually Means](https://www.buildseo.org/blog/cost-per-link-what-it-actually-means.md): Every link-building platform quotes a different number. Here's how to calculate the one that matters: cost per link still live at 90 days. And why that number is the only one your CFO should care about.
- [The Link Database Advantage: Why Campaign Two Should Cost 20% Less Than Campaign One](https://www.buildseo.org/blog/link-building-database-compounding.md): Campaign one teaches you everything about your vertical. Campaign two should cost less, take less time, and deliver more verified links â€” because you kept what you learned.

## See the Score Run on Your Category

We will qualify live publisher inventory in your vertical and show you every signal behind it. Schedule a demo: https://calendly.com/buildseo-sales/30min

---

[All posts](https://www.buildseo.org/blog.md) · [BuildSEO AI](https://www.buildseo.org/index.md)

